[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"detail-sidebar-cat-0-en-105":3,"doc-seo-128852-105":59,"doc-detail-128852-en":131},{"code":4,"msg":5,"data":6},0,"success",[7,13,18,23,28,33,38,43,48,51,55],{"id":8,"doc_module":4,"doc_module_name":9,"category_name":10,"show_sort_weight":11,"slug":12},1,"Document","Story & Novel",90,"story-novel",{"id":14,"doc_module":4,"doc_module_name":9,"category_name":15,"show_sort_weight":16,"slug":17},2,"Literature",80,"literature",{"id":19,"doc_module":4,"doc_module_name":9,"category_name":20,"show_sort_weight":21,"slug":22},4,"Exam",70,"exam",{"id":24,"doc_module":4,"doc_module_name":9,"category_name":25,"show_sort_weight":26,"slug":27},5,"Comic",60,"comic",{"id":29,"doc_module":4,"doc_module_name":9,"category_name":30,"show_sort_weight":31,"slug":32},6,"Technology",50,"technology",{"id":34,"doc_module":4,"doc_module_name":9,"category_name":35,"show_sort_weight":36,"slug":37},7,"Healthcare",40,"healthcare",{"id":39,"doc_module":4,"doc_module_name":9,"category_name":40,"show_sort_weight":41,"slug":42},8,"Research & Report",30,"research-report",{"id":44,"doc_module":4,"doc_module_name":9,"category_name":45,"show_sort_weight":46,"slug":47},9,"Religion & Spirituality",20,"religion-spirituality",{"id":46,"doc_module":4,"doc_module_name":9,"category_name":49,"show_sort_weight":46,"slug":50},"World Cup","world-cup",{"id":52,"doc_module":4,"doc_module_name":9,"category_name":53,"show_sort_weight":52,"slug":54},10,"Lifestyle","lifestyle",{"id":56,"doc_module":4,"doc_module_name":9,"category_name":57,"show_sort_weight":24,"slug":58},19,"General","general",{"code":4,"msg":60,"data":61},"ok",{"site_id":62,"language":63,"slug":64,"title":65,"keywords":66,"description":67,"schema_data":68,"social_meta":124,"head_meta":126,"extra_data":128,"updated_unix":130},105,"en","bridging-discrete-and-continuous-a-multimodal-strategy-for-complex-emotion-detection","BRIDGING DISCRETE AND CONTINUOUS - A MULTIMODAL STRATEGY FOR COMPLEX EMOTION DETECTION","","Human emotion recognition in human-computer interaction remains challenging because affective expressions are subtle, diverse, and strongly context-dependent. This study presents a multimodal framework that bridges discrete emotion categories and a continuous three-dimensional Valence-Arousal-Dominance (VAD) space, enabling smooth transitions between representation formats and generation of nuanced emotion labels. Discrete annotations are converted to VAD coordinates via K-means clustering, then a classifier predicts emotions in this joint space using facial expressions, vocal prosody, and textual transcripts. Evaluation on MER2024 shows preserved classification accuracy alongside improved interpretability and expressiveness compared with discrete-only approaches.",{"@graph":69,"@context":123},[70,84,106],{"@type":71,"itemListElement":72},"BreadcrumbList",[73,77,79,82],{"item":74,"name":75,"@type":76,"position":8},"https://docshare.wps.com","Home","ListItem",{"item":78,"name":9,"@type":76,"position":14},"https://docshare.wps.com/document/",{"item":80,"name":40,"@type":76,"position":81},"https://docshare.wps.com/document/research-report/",3,{"item":83,"name":65,"@type":76,"position":19},"https://docshare.wps.com/document/bridging-discrete-and-continuous-a-multimodal-strategy-for-complex-emotion-detection/128852/",{"url":83,"name":65,"@type":85,"image":86,"author":91,"headline":65,"publisher":94,"fileFormat":97,"inLanguage":63,"description":67,"dateModified":98,"datePublished":99,"encodingFormat":97,"isAccessibleForFree":100,"interactionStatistic":101},"DigitalDocument",{"url":87,"@type":88,"width":89,"height":90},"https://docshare.wps.com/thumbnails/bridging-discrete-and-continuous-a-multimodal-strategy-for-complex-emotion-detection/128852.png","ImageObject",300,407,{"name":92,"@type":93},"Aria","Person",{"url":74,"name":95,"@type":96},"DocShare","Organization","application/pdf","2026-09-23","2026-08-06",true,{"@type":102,"interactionType":103,"userInteractionCount":105},"InteractionCounter",{"@type":104},"ViewAction",15,{"@type":107,"mainEntity":108},"FAQPage",[109,115,119],{"name":110,"@type":111,"acceptedAnswer":112},"How does the framework bridge discrete emotions and continuous emotion representations?","Question",{"text":113,"@type":114},"It maps discrete emotion labels into a continuous Valence-Arousal-Dominance (VAD) space by converting annotations into VAD coordinates using K-means clustering, then classifies in that shared space for both discrete and continuous settings.","Answer",{"name":116,"@type":111,"acceptedAnswer":117},"Which modalities are used for emotion detection in the proposed system?",{"text":118,"@type":114},"The system integrates facial expressions, vocal prosody, and textual transcripts from video clips to predict Valence-Arousal-Dominance scores.",{"name":120,"@type":111,"acceptedAnswer":121},"How is the method evaluated and what dataset is used?",{"text":122,"@type":114},"The framework is evaluated on the MER2024 dataset, which provides culturally consistent Chinese media video clips annotated with both discrete and open-vocabulary emotion labels, enabling assessment across different emotion representation formats.","https://schema.org",{"og:url":83,"og:type":125,"og:title":65,"og:site_name":95,"og:description":67},"article",{"robots":127,"canonical":83},"index,follow",{"doc_id":129,"site_id":62},128852,1786003892,{"code":4,"msg":5,"data":132},{"doc_id":129,"user_id":133,"nickname":92,"user_avatar":134,"doc_module":4,"category_id":39,"category_name":40,"doc_title":65,"doc_description":67,"doc_content":135,"file_id":136,"file_url":137,"file_type":138,"file_size":139,"view_count":105,"is_deleted":4,"is_public":8,"is_downloadable":8,"audit_status":8,"page_count":29,"language":140,"language_code":63,"site_id":62,"html_lang":63,"table_of_contents":141,"faqs":142,"seo_title":143,"seo_description":67,"update_tm":130,"read_time":105},2336474459895,"https://ap-avatar.wpscdn.com/avatar/22000baeef7a5ed0655?x-image-process=image/resize,m_fixed,w_180,h_180&k=1786071322749376916","BRIDGING DISCRETE AND CONTINUOUS: A MULTIMODAL STRATEGY FOR COMPLEX EMOTION DETECTION  \nJiehui Jia Huan Zhang Jinhua Liang Queen Mary University of London, Centre for Digital Music, London, United Kingdom  \nABSTRACT  \nRecognizing and interpreting human emotions remains a core challenge in human-computer interaction, due to the subtlety, diversity, and context-dependence of affective expressions. This study proposesa novel framework for mapping emotions between discrete categories and a continuous three-dimensional Valence-Arousal-Dominance (VAD) space, which allows for seamless transitions between discrete and continuous emotion representations and supports the generation of nuanced, semantically rich emotion labels. Using K-means clustering, we convert traditional discrete annotations into continuous coordinates in VAD space and build a classifier that operates within this framework. The system integrates a multimodal approach, combining facial expressions, vocal prosody, and textual transcripts from video clips to detect emotions across a broad and flexible range. The framework is evaluated on the MER2024 dataset, which contains culturally consistent video clips from Chinese media annotated with both discrete and open-vocabulary emotion labels, providing a robust testbed for assessing performance across different emotion representation formats. Our method demonstrates that the mapping framework preserves classification accuracy while significantly enhancing the model’s ability to reflect the complexity and variability of real human emotions, enabling more interpretable and expressive emotion recognition compared to traditional discrete-only methods.  \nIndex Terms— Multimodal emotion recognition, emotional variability, valence-arousal-dominance (VAD) framework, emotion detection, machine learning  \n1. INTRODUCTION  \nHuman emotions are complex and are described through diverse vocabularies across languages, reflecting our thoughts, feelings, and reactions via facial expressions, body language, voice tone, and speech [1] . Accurate comprehension and response to human emotions by machines can significantly benefit areas such as marketing, mental health monitoring, multimedia generation, and human-computer interaction [2, 3, 4, 5, 6, 7] . Therefore, developing systems that can accurately recognize the variety of human emotions is essential.  \nHowever, the challenge in emotion detection lies in the subjective nature of emotions. It is hard to set a clear boundary to categorize emotions, so as to choose a “basic” emotion group [8] . Moreover, emotion datasets vary in their annotation schemes (e.g., differing discrete labels) and domains, hindering direct comparisons across previous works. These variations restrict prior research to specific data sources, limiting their generalizability to real-world applications [9] .  \nRecognizing the challenges, this paper proposes a multimodal framework that transforms different discrete emotion labels into one continuous emotion label framework. We propose to use a fixed  \nFig. 1. Emotion Vocabularies in 3D VAD Space.  \nThe 3D VAD emotion space is based on the foundational work by Russell (1980) [10], and the emotion lexicon used for mapping is derived from the NRC-VAD lexicon developed by Mohammad (2018) [11]  \nemotion Valence, Arousal, and Dominance rating scale to standardize the scoring of a variety of emotion labels in a multidimensional space. Using K-means clustering, we build a classifier that transitions between discrete and continuous emotions, supporting both closedset and open-set emotion recognition tasks. In our case, we grouped the emotion labels into six clusters based on the six basic emotion labels which have been annotated for the dataset we chose. We built a multimodal model that integrates facial expressions, voice tones, and transcripts from video clips to predict the Valence-ArousalDominance (VAD) scores of the emotions. Finally, the VAD score was mapped back to the origi","cbCaitmoJ5dq2UVX","https://ap.wps.com/l/cbCaitmoJ5dq2UVX","pdf",1106598,"English","# Introduction\n## Motivation and challenges in emotion detection\n## Proposed multimodal discrete-to-continuous framework\n# Related work\n## Measurement of emotion","[{\"question\":\"How does the framework bridge discrete emotions and continuous emotion representations?\",\"answer\":\"It maps discrete emotion labels into a continuous Valence-Arousal-Dominance (VAD) space by converting annotations into VAD coordinates using K-means clustering, then classifies in that shared space for both discrete and continuous settings.\"},{\"question\":\"Which modalities are used for emotion detection in the proposed system?\",\"answer\":\"The system integrates facial expressions, vocal prosody, and textual transcripts from video clips to predict Valence-Arousal-Dominance scores.\"},{\"question\":\"How is the method evaluated and what dataset is used?\",\"answer\":\"The framework is evaluated on the MER2024 dataset, which provides culturally consistent Chinese media video clips annotated with both discrete and open-vocabulary emotion labels, enabling assessment across different emotion representation formats.\"}]","BRIDGING DISCRETE AND CONTINUOUS - A MULTIMODAL STRATEGY FOR COMPLEX EMOTION DETECTION | PDF"]