[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-134748-en":3,"doc-seo-134748-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},134748,687197207639,"Asher","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","SEDA - Simple and Effective Data Augmentation for Sign Language Understanding","Sign language understanding (SLU) converts sign language videos into glosses and produces corresponding outputs, covering sign language recognition (SLR) and sign language translation (SLT). This task is difficult due to fine-grained video understanding and sequence generation, compounded by limited supervised training data. To narrow the vision-language modality gap and address data scarcity, SEDA introduces simple, effective data augmentation on both sign and text sides plus multi-task learning with task-specific fine-tuning. Experiments on RWTH-PHOENIX Weather 2014T show consistent gains with WER 19.91, BLEU 25.19, and ROUGE 51.72.","SEDA: Simple and Effective Data Augmentation for Sign Language  \nUnderstanding  \nSihan Tan 1 ,2 , Taro Miyazaki2 , Katsutoshi Itoyama 1 ,3 , Kazuhiro Nakadai 1   \n1Tokyo Institute of Technology,  \n2 NHK Science and Technology Research Laboratories,  \n3 Honda Research Institute Japan Co. , Ltd.  \n{tansihan, itoyama, [nakadai}@ra.sc.e.titech.ac.jp](nakadai}@ra.sc.e.titech.ac.jp), [miyazaki.t-jw@nhk.or.jp](miyazaki.t-jw@nhk.or.jp)  \nAbstract  \nSign language understanding (SLU) aims to convert sign language videos into glosses that transcribe sign language word-by-word by means of another written language and generate corresponding spoken sentences, including sign language recognition (SLR) and sign language translation (SLT) . SLU has been a challenging undertaking since it demands the capability of fine-grained video understanding and sequence generation. In addition, the lack of supervised training data further hinders the advancement of SLU. To narrow the modality gap between vision and language and mitigate the data scarcity problem, we propose a Simple and Effective Data Augmentation (SEDA) framework for end-to-end SLU. In particular, SEDA consists of two key components: data augmentations on both sign and text sides and multi-task learning with task-specific fine-tuning. Experimental results on RWTH-PHOENIX Weather 2014T demonstrate that our proposed SEDA framework significantly and consistently outperforms the baseline model and achieves a WER of 19.91, a BLEU score of 25.19, and a ROUGE score of 51.72, delivering competitive scores in both SLR and SLT.  \nKeywords: Sign language understanding, Data augmentation, Multi-task learning.  \n1. Introduction  \nAs the native language used by deaf and hardof-hearing individuals to communicate, sign languages (SLs) exhibit distinctive grammar and have been established as a form of natural language (Klima and Bellugi , 1979) . Sign language understanding (SLU) in which SLs are understood by means of machines mainly involves two functions: sign language recognition (SLR) and sign language translation (SLT) . It is a challenging undertaking that requires the model to have the capability of fine-grained video understanding and sequence generation. Unlike spoken languages, SLs involve manual and non-manual elements (e.g., the movement of the body, head, mouth, or even eyebrows) . Also, the visual signal in SLs displays dramatic variability among signers, posing a huge modality gap when transforming SLs into text (Zhang et al. , 2023) . Insufficient supervised training data presents an additional challenge to the advancement of SLU, as it increases the risk of overfitting. To tackle these challenges, it is essential to devise inductive biases, such as novel model architectures, training strategies and objectives, facilitating knowledge transfer, and the induction of universal representations for SLU. In this paper, we aim to augment SLs data on both sign and text sides, and provide effective training, including multi-task learning.  \nExisting SLU methods follow the framework of  \nneural machine translation (NMT) where the source language is spatial-temporal pixels rather than discrete tokens and the target language is spoken languages. Depending on the model architectures, annotation pairs, or final goals, SLU comprises: Sign2Gloss (Min et al. , 2021 ; Hao et al. , 2021), Sign2Gloss2Text (Yin and Read , 2020), Sign2(Gloss+Text) (Camgoz et al. , 2020) and Sign2Text (Camgoz et al. , 2018 ; Chen et al. , 2022) tasks. Additionally, to boost the well-being of the sign language community and improve SLU performance, a number of studies have focused on Gloss2Text (Moryossef et al. , 2021) and Text2Gloss (Miyazaki et al. , 2020 ; Zhu et al. , 2023) by transfer learning, data augmentation, etc. Following this line of study, we find that researchers seldom explore data augmentation techniques for the sign aspect, primarily concentrating on the textual component. Furthermore, constructing large-scale","cbCaifIa4RHKyxiJ","https://ap.wps.com/l/cbCaifIa4RHKyxiJ","pdf",508010,1,6,"English","en",105,"# Introduction\n## Sign language understanding overview\n## Challenges in SLU (modality gap and data scarcity)\n# Related Work\n## Sign language understanding approaches (cascading vs end-to-end)","[{\"question\":\"What does SLU aim to achieve in sign language understanding?\",\"answer\":\"SLU converts sign language videos into glosses that transcribe sign language word-by-word and can generate corresponding spoken sentences. It covers sign language recognition (SLR) and sign language translation (SLT).\"},{\"question\":\"Why is SLU challenging according to the document?\",\"answer\":\"SLU requires fine-grained video understanding and sequence generation, and suffers from a lack of supervised training data. Sign languages also show strong variability in manual and non-manual elements, creating a modality gap when converting to text.\"},{\"question\":\"How does SEDA improve end-to-end SLU performance?\",\"answer\":\"SEDA applies data augmentations to both sign and text inputs and uses multi-task learning with task-specific fine-tuning. The method increases training samples and helps the model learn highly related tasks.\"}]","SEDA - Simple and Effective Data Augmentation for Sign Language Understanding | PDF",1787298999,15,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"seda-simple-and-effective-data-augmentation-for-sign-language-understanding","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/seda-simple-and-effective-data-augmentation-for-sign-language-understanding/134748/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-22","2026-08-21",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What does SLU aim to achieve in sign language understanding?","Question",{"text":76,"@type":77},"SLU converts sign language videos into glosses that transcribe sign language word-by-word and can generate corresponding spoken sentences. It covers sign language recognition (SLR) and sign language translation (SLT).","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"Why is SLU challenging according to the document?",{"text":81,"@type":77},"SLU requires fine-grained video understanding and sequence generation, and suffers from a lack of supervised training data. Sign languages also show strong variability in manual and non-manual elements, creating a modality gap when converting to text.",{"name":83,"@type":74,"acceptedAnswer":84},"How does SEDA improve end-to-end SLU performance?",{"text":85,"@type":77},"SEDA applies data augmentations to both sign and text inputs and uses multi-task learning with task-specific fine-tuning. The method increases training samples and helps the model learn highly related tasks.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":93},[94,98,102,106,111,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":21,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":107,"slug":138},19,"General","general"]