[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82572-en":3,"doc-seo-82572-105":30,"detail-sidebar-cat-0-en-105":90},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},82572,962075006959,"Anda","https://ap-avatar.wpscdn.com/avatar/e0002397efbe92a78e?_k=1776741047341049297",8,"Research & Report","Quantifying the Affective Gap A Zero-Shot Evaluation of LLMs on Fine-Grained Emotion Taxonomies","Emotion recognition in natural language is a core challenge in affective computing, with direct impact on human-computer interaction, mental health support, and conversational AI. The study provides a unified zero-shot API evaluation of three commercial LLMs—Claude, ChatGPT, and Gemini—on a 13-class fine-grained emotion classification task using a stratified 1,000-sentence sample from the boltuix/emotions dataset. Gemini leads with 39.9% accuracy and 0.363 macro-F1, followed by GPT-5.4 and Claude, which shows strong macro-F1 bias due to class imbalance. All models largely fail on love, confusion, and shame.","Quantifying the Affective Gap: A Zero-Shot Evaluation of LLMs on Fine-Grained Emotion  \nTaxonomies  \nLawrence Obiuwevwi, Krzysztof J. Rechowicz, Jessica M. Johnson, Vikas Ashok, Sachin Shetty, & Sampath Jayarathna  \nOld Dominion University, Norfolk, VA, USA  \n[lobiu001@odu.edu](lobiu001@odu.edu), [krechowi@odu.edu](krechowi@odu.edu), [j17johnso@odu.edu](j17johnso@odu.edu),  \n[vganjiqu@cs.odu.edu](vganjiqu@cs.odu.edu), [sshetty@odu.edu](sshetty@odu.edu), [sampath@cs.odu.edu](sampath@cs.odu.edu)  \narXiv :2607 .00968v 1 [ cs .CL] 1 Jul 2026  \nAbstract—Emotion recognition in natural language is a foundational challenge in affective computing, with critical implications for human-computer interaction, mental health support, and conversational AI. This paper presents a rigorous, unified zeroshot evaluation of three leading commercial large language models: Claude (claude-sonnet-4-6), ChatGPT (GPT-5.4), and Gemini (gemini-2.5-flash). The models were queried through their respective production APIs as of April 2026 on a fine-grained 13-class emotion classification task. Using a stratified 1,000-sentence sample from the boltuix/emotions-dataset comprising 131,306 sentences across 13 categories, a single uniform prompt with no exemplars was applied identically across all models. Gemini achieves the highest accuracy (39.9%) and macro-F1 (0.363), followed by GPT-5.4 (38.8%, F1 = 0.291) and Claude (38.0%, F1 = 0.159). All models excel on sarcasm and desire while consistently failing on love, confusion, and shame. McNemar’s tests reveal no statistically significant pairwise differences (p > 0. 10), suggesting convergence at a shared zero-shot ceiling. Claude’s markedly lower macro-F1 exposes a class-imbalance prediction bias. These findings highlight the current limitations of frontier AI systems in zero-shot fine-grained emotion classification.  \nIndex Terms—emotion classification, affective computing, LLMs, zero-shot learning, NLP, GPT-5.4, Claude, Gemini, AI, human-computer interaction  \nI. INTRODUCTION  \nEmotions shape human communication and underpin highstakes applications including mental health monitoring [1], empathetic dialogue systems [2], and human-computer interaction. The ability of an AI system to recognize affective states is therefore a prerequisite for deployment in any domain where human welfare is at stake.  \nDespite rapid advances in large language models (LLMs), their affective intelligence remains poorly characterized at fine-grained resolution. Most studies probe coarse polarity or the Ekman six basic emotions [3] rather than richer taxonomies [4], [5], and direct cross-provider comparisons under identical experimental conditions are absent from the literature. This gap matters: practitioners selecting an API for emotion-aware applications must rely on ad hoc or proprietary benchmarks that may not reflect real-world linguistic diversity. This paper addresses the gap through four research questions: zero-shot accuracy across 13 classes (RQ1); pairwise  \nstatistical differences using McNemar’s test (RQ2); per-class strengths and failure modes (RQ3); and the effect of sentence length on accuracy (RQ4) . We contribute (i) the first direct zero-shot comparison of Claude (claude-sonnet-4-6), ChatGPT (GPT-5.4), and Gemini (gemini-2.5-flash) on a 13-class task using production APIs as of April 2026; (ii) per-emotion accuracy breakdowns revealing systematic failure modes; (iii) McNemar significance testing; and (iv) a sentence-length moderation analysis.  \nII. RELATED WORK  \nA. Emotion Classification and Datasets  \nComputational emotion recognition has progressed from lexicon-based resources such as the NRC Lexicon [6] and SemEval affective tasks [7]–[9] through deep learning classifiersto transformer fine-tuning paradigms achieving state-of-theart results on benchmarks including GoEmotions [10](58K comments, 27 categories) . BERT [11],[12] and RoBERTa [13] established the transformer paradigm as dominant for supervised emoti","cbCaidDbS7YkbnYg","https://ap.wps.com/l/cbCaidDbS7YkbnYg","pdf",478624,2,1,4,"English","en",105,"# Introduction\n# Related Work\n## Emotion Classification and Datasets\n## LLMs for Emotion and Sentiment\n# Methodology\n## Dataset","[{\"question\":\"What is the main research goal of the paper?\",\"answer\":\"The paper quantifies how well frontier LLMs perform in zero-shot fine-grained emotion classification across 13 emotion classes under identical experimental conditions.\"},{\"question\":\"Which LLMs are evaluated and how are they queried?\",\"answer\":\"Claude, ChatGPT (GPT-5.4), and Gemini are queried through their production APIs using a single uniform prompt without exemplars.\"},{\"question\":\"What dataset and evaluation setup are used for the 13-class task?\",\"answer\":\"A stratified random sample of 1,000 English sentences is drawn from the boltuix/emotions dataset (131,306 labeled sentences across 13 mutually exclusive emotions), preserving class proportions for realistic deployment simulation.\"}]",1784181588,10,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":85,"head_meta":87,"extra_data":89,"updated_unix":28},"quantifying-the-affective-gap-a-zero-shot-evaluation-of-llms-on-fine-grained-emotion-taxonomies","",{"@graph":36,"@context":84},[37,52,67],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":22},"https://docshare.wps.com/document/quantifying-the-affective-gap-a-zero-shot-evaluation-of-llms-on-fine-grained-emotion-taxonomies/82572/",{"url":51,"name":13,"@type":53,"author":54,"headline":13,"publisher":56,"fileFormat":59,"inLanguage":24,"description":14,"dateModified":60,"datePublished":61,"encodingFormat":59,"isAccessibleForFree":62,"interactionStatistic":63},"DigitalDocument",{"name":9,"@type":55},"Person",{"url":41,"name":57,"@type":58},"DocShare","Organization","application/pdf","2026-07-22","2026-07-16",true,{"@type":64,"interactionType":65,"userInteractionCount":20},"InteractionCounter",{"@type":66},"ViewAction",{"@type":68,"mainEntity":69},"FAQPage",[70,76,80],{"name":71,"@type":72,"acceptedAnswer":73},"What is the main research goal of the paper?","Question",{"text":74,"@type":75},"The paper quantifies how well frontier LLMs perform in zero-shot fine-grained emotion classification across 13 emotion classes under identical experimental conditions.","Answer",{"name":77,"@type":72,"acceptedAnswer":78},"Which LLMs are evaluated and how are they queried?",{"text":79,"@type":75},"Claude, ChatGPT (GPT-5.4), and Gemini are queried through their production APIs using a single uniform prompt without exemplars.",{"name":81,"@type":72,"acceptedAnswer":82},"What dataset and evaluation setup are used for the 13-class task?",{"text":83,"@type":75},"A stratified random sample of 1,000 English sentences is drawn from the boltuix/emotions dataset (131,306 labeled sentences across 13 mutually exclusive emotions), preserving class proportions for realistic deployment simulation.","https://schema.org",{"og:url":51,"og:type":86,"og:title":13,"og:site_name":57,"og:description":14},"article",{"robots":88,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":91},[92,96,100,104,109,114,119,122,127,130,133],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":93,"show_sort_weight":94,"slug":95},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":97,"show_sort_weight":98,"slug":99},"Literature",80,"literature",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":101,"show_sort_weight":102,"slug":103},"Exam",70,"exam",{"id":105,"doc_module":4,"doc_module_name":46,"category_name":106,"show_sort_weight":107,"slug":108},5,"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":46,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":29,"doc_module":4,"doc_module_name":46,"category_name":131,"show_sort_weight":29,"slug":132},"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":46,"category_name":135,"show_sort_weight":105,"slug":136},19,"General","general"]