[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86407-en":3,"doc-seo-86407-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86407,3848291630094,"Emma Wilson","https://eur-avatar.wpscdn.com/davatar_085a072bc5b1113ac321206ff7593b45",8,"Research & Report","Emotion Recognition in Sign Language Conversation","Emotion recognition in conversation is central to affective computing, yet sign language emotion datasets largely emphasize isolated sentences and omit conversational context, causing real-world performance degradation. To overcome this structural gap, the document introduces the ERC task for sign language video analysis and proposes the eJSL Dialog dataset. Built from STUDIES dialogue scripts, it provides 1,920 video samples across 480 unique teacher–student dialogues. Benchmarks across visual, text-based, and multimodal models show a domain gap and highlight the need for context-aware sign-language visual extractors and larger datasets for scalable pre-training.","Emotion Recognition in Sign Language Conversation  \nYusong Wang 1 , Keyu Mao 1 , Takao Obi 1 , Minghao Shao2 and Kotaro Funakoshi 1  \narXiv :2605 .23328v2 [ cs .CL] 13 Jul 2026  \nAbstract—Emotion Recognition in Conversation is a core component of affective computing, while current sign language emotion datasets primarily focus on isolated sentences and lack conversational context. Models trained exclusively on these isolated utterances demonstrate degraded performance in real world scenarios because they cannot utilize historical dialogue flow. To address this structural limitation, we introduce the ERC task to sign language video analysis and propose theeJSL Dialog dataset. Constructed using the scripts from the STUDIES corpus, the dataset contains 1,920 video samples organized into 480 unique dialogues. We conduct systematic benchmarking on this dataset using models ranging from isolated visual networks to multimodal conversational architectures. The results reveal a domain gap when applying generic multimodal conversational emotion recognition models to sign language. These findings demonstrate the explicit need for context-aware visual extractors specific to sign language and indicate that constructing larger conversational datasets to support large-scale pre-training is a necessary next step for future research.  \nI. INTRODUCTION  \nEmotion recognition enables machines to perceive users’affective states and provide empathetic responses, making it central to virtual assistants and emotionally aware assistive technologies [1], [2], [3] . Most systems focus on spoken languages and standard facial expressions from non-signing populations. Extending these technologies to sign language, which is often ignored in mainstream research, is a necessary step [4], [5] . In sign language communication, the emotion recognition task becomes highly complex due to the visual nature of the language itself. As visual languages, sign languages rely on hand signs, facial expressions, and upper body movements to convey linguistic structure and emotional content simultaneously [6], [7] . This overlap between grammatical features and emotional features introduces significant ambiguity to automatic emotion recognition systems [6], [8] . Accurately capturing and modeling emotional dynamics in sign language remains a practical challenge in the field of both computer vision and language processing.  \nExisting resources for sign language emotion analysis focus primarily on isolated sentences or unidirectional expressions. For instance, datasets such as eJSL Solo [9] consist of sign language video clips detached from conversational context. Similarly, the EmoSign dataset [10] concentrates on capturing emotional expressions within single video utterances. These datasets have advanced research in isolated sign  \n1 Yusong Wang, Keyu Mao, Takao Obi, and Kotaro Funakoshi are with Institute of Science Tokyo, Japan. {wangyi, maokeyu, obi, [funakoshi](funakoshi}@lr.first.iir.isct.ac.jp)[}](funakoshi}@lr.first.iir.isct.ac.jp)[@lr.first.iir.isct.ac.jp](funakoshi}@lr.first.iir.isct.ac.jp)  \n2 Minghao Shao is with New York University, New York, 11201, USA. [shao.minghao@nyu.edu](shao.minghao@nyu.edu)  \nTABLE I  \nEXAMPLE OF DIALOGUE LINES FROM STUDIES [14] BY A TEACHER AND A MALE STUDENT USED IN OUR EJSL DIALOG DATASET.  \n\n| Speaker Emotion Line |  |  |\n| --- | --- | --- |\n| Male student | Joy | 先生！この前部活の試合で勝ったんだ!(Teacher! I won my club match the other day!) |\n| Teacher | Joy | 文武両道だね!\u003Cbr>(You are excelling in both academics and sports!) |\n| Male student | Joy | そう!それを目指してる!\u003Cbr>(Yes! That is what I am aiming for!) |\n| Teacher | Neutral | あなたなら出来るわ。これからもしっかり頑張るのよ!\u003Cbr>(You can do it. Keep working hard from now on!) |\n\nlanguage emotion recognition, but they omit the conversational context present in real communication.  \nIn general, emotion recognition models trained exclusively on these isolated utterances cannot utilize historical context. Consequently, they degrade ","cbCaitjah0BJnHjN","https://ap.wps.com/l/cbCaitjah0BJnHjN","pdf",3033707,4,1,7,"English","en",105,"# Introduction\n## Task Motivation and Challenges\n## Limitations of Existing Datasets\n## Proposed ERC Task and eJSL Dialog Dataset\n## Benchmarking and Baseline Results","[{\"question\":\"Why does emotion recognition perform worse in real scenarios for sign language models trained on isolated utterances?\",\"answer\":\"Isolated-utterance training cannot leverage historical dialogue flow, so the models miss how emotional meaning depends on continuous turn-by-turn context.\"},{\"question\":\"What is the eJSL Dialog dataset and how is it constructed?\",\"answer\":\"The dataset is built from dialogue scripts from the STUDIES Japanese Empathetic Dialogue Speech Corpus and organizes 1,920 video samples into 480 unique teacher–student dialogues with emotion labels per line.\"},{\"question\":\"What do benchmark experiments show about applying generic multimodal conversational emotion recognition to sign language?\",\"answer\":\"Benchmarking indicates a domain gap: generic multimodal conversational models fail to capture sign-language-specific emotional dynamics, especially when visual extractors lack context-aware capability.\"}]",1784211551,18,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"emotion-recognition-in-sign-language-conversation","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/emotion-recognition-in-sign-language-conversation/86407/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-28","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why does emotion recognition perform worse in real scenarios for sign language models trained on isolated utterances?","Question",{"text":75,"@type":76},"Isolated-utterance training cannot leverage historical dialogue flow, so the models miss how emotional meaning depends on continuous turn-by-turn context.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What is the eJSL Dialog dataset and how is it constructed?",{"text":80,"@type":76},"The dataset is built from dialogue scripts from the STUDIES Japanese Empathetic Dialogue Speech Corpus and organizes 1,920 video samples into 480 unique teacher–student dialogues with emotion labels per line.",{"name":82,"@type":73,"acceptedAnswer":83},"What do benchmark experiments show about applying generic multimodal conversational emotion recognition to sign language?",{"text":84,"@type":76},"Benchmarking indicates a domain gap: generic multimodal conversational models fail to capture sign-language-specific emotional dynamics, especially when visual extractors lack context-aware capability.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,119,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]