[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85187-en":3,"doc-seo-85187-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},85187,962075114101,"Seraphina","https://ap-avatar.wpscdn.com/avatar/e000253a75eb197efd?x-image-process=image/resize,m_fixed,w_180,h_180&k=1780044092746381165",8,"Research & Report","Benchmarking Dynamic Affective Reasoning Viewer-Centric Video Emotion Dataset DAR","Video emotion analysis is commonly treated as static clip-level classification, ignoring how emotions evolve through cumulative reactions to consecutive causal events. DAR (Dynamic Affective Reasoning) introduces a large-scale viewer-centric benchmark for affect transitions and causal reasoning across temporal video events. DAR includes 15,087 videos with 36,908 event-aligned affective segments annotated across 27 emotion categories, offering dense, temporally grounded, causally explicit reasoning chains. Using DAR, three tasks are defined: affective segmentation, fine-grained emotion classification, and affective reasoning, and a two-stage DAR-R1 model is proposed and evaluated on 10+ MLLMs.","arXiv :2607 . 10238v1 [ cs .CV] 11 Jul 2026  \nBenchmarking Dynamic Affective Reasoning: A Viewer-Centric Video Emotion Dataset  \nZhiyan Zhang 1, Peipei Song 1⋆ , Jinpeng Hu2, Jingyang Jia 1, Xun Yang 1, and Xiaojun Chang 1  \n1 University of Science and Technology of China, Hefei, China [zzyhang02@gmail.com](zzyhang02@gmail.com) , [beta.songpp@gmail.com](beta.songpp@gmail.com) , [jjygood@mail.ustc.edu.cn](jjygood@mail.ustc.edu.cn) ,  \n[xyang21@ustc.edu.cn](xyang21@ustc.edu.cn) , [xjchang@ustc.edu.cn](xjchang@ustc.edu.cn)  \n2 Hefei University of Technology, Hefei, China  \n[135858hjp@gmail.com](135858hjp@gmail.com)  \nAbstract. Video emotion analysis is typically framed as a static classification problem, treating each clip as an independent labeled unit.  \nHowever, such a formulation overlooks a key psychological fact: emotions change as a result of cumulative reactions to consecutive causal events. To bridge this gap, we introduce DAR (Dynamic Affective Reasoning), the first large-scale benchmark for viewer-centric affect transitions and causal reasoning over consecutive video events. DAR contains 15,087 videos and  \n36,908 event-aligned affective segments annotated with 27 emotion categories. Unlike existing video-based emotion datasets, DAR presents anew viewer-centric perspective on fine-grained emotional expressions and transitions, and provides dense, temporally grounded, and causally explicit reasoning chains. Based on DAR, we formally define three challenging tasks: affective segmentation, fine-grained emotion classification, and affective reasoning. Complementing this benchmark, we propose DARR1 , a two-stage framework that combines supervised fine-tuning with Group Relative Policy Optimization. Experiments across 10+ MLLMs show that DAR-R1 sets a new state-of-the-art for dynamic affective reasoning, in terms of both emotional localization and affective reasoning.  \nProject page: [https://github.com/Zhang-Zhiyan/DAR](https://github.com/Zhang-Zhiyan/DAR).  \nKeywords: Dynamic Affective Reasoning · Video Emotion Analysis · Large-scale Dataset · Reinforcement Learning  \n1 Introduction  \nAffective computing [31] aims to equip machines with the ability to perceive, interpret, and reason about human affect, enabling applications from human– computer interaction [13] to psychological counseling and psychotherapy [27,43, 44,48] . As AI systems become embedded in everyday products [45], robust emo  \ntion understanding is increasingly essential for natural, adaptive, and socially aware interaction [37] . Recent Multimodal Large Language Models (MLLMs) [1,⋆ Corresponding author.  \n2 Z. Zhang et al.  \n0.0s 5.5s 8.5s 16. 1s  \nTask 1: Affective Segmentation Split 1 Split 2  \nTemporal Localization  0.0s-5 .5s  5.5s-8 .5s  8.5-16. 1s   \nTask 2: Fine-grained Emotion Classification  \nFrom 27 Categories  Anxiety  Fear  Empathic Pain   \n0.0s 5.5s 8.5s 16. 1s  \n[{Start: 0.0s, End: 5.5s. },  \n{Start: 5.5s, End: 8.5s. },  \n{Start: 8.5s, End: 16.1s. }]  \n[{Start: 0.0s, End: 5.5s, Emotion: Anxiety. },{Start: 5.5s, End: 8.5s, Emotion: Fear. },  \n{Start: 8.5s, End: 16.1s, Emotion: Empathic Pain. }]  \nTask 3: Affective Reasoning [{0 .0s-5.5s, Anxiety, The viewer feels anxiety due to the stark …},  \nCausal Logic Explanation 0.0s ue tAnxioet…y 5.B5s auseFear 8. .5s EThempamthain ic Psa 16. 1s{5{8 ..5s5s--8.165s1,searEm,ThepathivciewPaerin,’sTehemvoitionewershnoiftswffromeelsaEnmxpietyathtoicPfeaainr},  \nFig. 1: Overview of the DAR Benchmark Tasks. We formulate three hierarchical tasks: (1) Affective Segmentation for locating temporal boundaries of emotion shifts;  \n(2) Fine-grained Emotion Classification utilizing a 27-category viewer-centric taxonomy; and (3) Affective Reasoning for generating causal explanations grounded in visual evidence.  \n9, 46, 47] have further advanced video-based emotion understanding and recognition [10, 16] . Foundational benchmarks such as DFEW [19] and VCE [29] enabled systematic evaluation of video emotion understanding,","cbCaihdKlzjQbJvY","https://ap.wps.com/l/cbCaihdKlzjQbJvY","pdf",2254709,1,25,"English","en",105,"# Abstract\n# Introduction\n## Dynamic Affective Reasoning Tasks\n## Benchmark Contributions and Annotation Design\n# Example Tasks and Outputs\n## Affective Segmentation\n## Fine-grained Emotion Classification\n## Affective Reasoning","[{\"question\":\"What problem does DAR address in existing video emotion analysis benchmarks?\",\"answer\":\"Existing benchmarks often treat video emotion as a static label for an entire clip. DAR targets the overlooked dynamic nature of emotions as they change due to consecutive causal events.\"},{\"question\":\"What does the DAR dataset contain and how is it annotated?\",\"answer\":\"DAR provides 15,087 videos and 36,908 event-aligned affective segments. Each segment is annotated with 27 emotion categories and aligned to event-level temporal segments.\"},{\"question\":\"What are the three benchmark tasks introduced with DAR?\",\"answer\":\"DAR defines affective segmentation to locate emotion-shift boundaries, fine-grained emotion classification using a 27-category taxonomy, and affective reasoning to generate causal logic grounded in visual evidence.\"}]",1784201625,63,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"benchmarking-dynamic-affective-reasoning-viewer-centric-video-emotion-dataset-dar","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/benchmarking-dynamic-affective-reasoning-viewer-centric-video-emotion-dataset-dar/85187/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does DAR address in existing video emotion analysis benchmarks?","Question",{"text":75,"@type":76},"Existing benchmarks often treat video emotion as a static label for an entire clip. DAR targets the overlooked dynamic nature of emotions as they change due to consecutive causal events.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What does the DAR dataset contain and how is it annotated?",{"text":80,"@type":76},"DAR provides 15,087 videos and 36,908 event-aligned affective segments. Each segment is annotated with 27 emotion categories and aligned to event-level temporal segments.",{"name":82,"@type":73,"acceptedAnswer":83},"What are the three benchmark tasks introduced with DAR?",{"text":84,"@type":76},"DAR defines affective segmentation to locate emotion-shift boundaries, fine-grained emotion classification using a 27-category taxonomy, and affective reasoning to generate causal logic grounded in visual evidence.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]