[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-seo-138300-105":3,"detail-sidebar-cat-0-en-105":81,"doc-detail-138300-en":130},{"code":4,"msg":5,"data":6},0,"ok",{"site_id":7,"language":8,"slug":9,"title":10,"keywords":11,"description":12,"schema_data":13,"social_meta":74,"head_meta":76,"extra_data":78,"updated_unix":80},105,"en","changing-the-representation-examining-language-representation-for-neural-sign-language-production","Changing the Representation - Examining Language Representation for Neural Sign Language Production","","Neural Sign Language Production (SLP) automatically converts spoken-language sentences into sign language videos. This work improves the first Text-to-Gloss step by applying NLP methods that enhance sentence-level embeddings using language models such as BERT and Word2Vec, along with multiple tokenization strategies for the low-resource Text to Gloss setting. The paper introduces Text to HamNoSys (T2H) translation and demonstrates benefits of phonetic HamNoSys over gloss-level representations. It further uses hand-shape extraction from HamNoSys as additional training supervision, achieving strong BLEU-4 results on MineDGS and PHOENIX14T.",{"@graph":14,"@context":73},[15,34,56],{"@type":16,"itemListElement":17},"BreadcrumbList",[18,23,27,31],{"item":19,"name":20,"@type":21,"position":22},"https://docshare.wps.com","Home","ListItem",1,{"item":24,"name":25,"@type":21,"position":26},"https://docshare.wps.com/document/","Document",2,{"item":28,"name":29,"@type":21,"position":30},"https://docshare.wps.com/document/research-report/","Research & Report",3,{"item":32,"name":10,"@type":21,"position":33},"https://docshare.wps.com/document/changing-the-representation-examining-language-representation-for-neural-sign-language-production/138300/",4,{"url":32,"name":10,"@type":35,"image":36,"author":41,"headline":10,"publisher":44,"fileFormat":47,"inLanguage":8,"description":12,"dateModified":48,"datePublished":49,"encodingFormat":47,"isAccessibleForFree":50,"interactionStatistic":51},"DigitalDocument",{"url":37,"@type":38,"width":39,"height":40},"https://docshare.wps.com/thumbnails/changing-the-representation-examining-language-representation-for-neural-sign-language-production/138300.png","ImageObject",300,407,{"name":42,"@type":43},"Levi","Person",{"url":19,"name":45,"@type":46},"DocShare","Organization","application/pdf","2026-09-18","2026-08-23",true,{"@type":52,"interactionType":53,"userInteractionCount":55},"InteractionCounter",{"@type":54},"ViewAction",9,{"@type":57,"mainEntity":58},"FAQPage",[59,65,69],{"name":60,"@type":61,"acceptedAnswer":62},"What does Neural Sign Language Production (SLP) aim to do?","Question",{"text":63,"@type":64},"SLP converts spoken-language sentences into sign language videos automatically. The task has traditionally been split into Text-to-Gloss translation and then video production from gloss sequences.","Answer",{"name":66,"@type":61,"acceptedAnswer":67},"How does this paper improve the Text-to-Gloss step in SLP?",{"text":68,"@type":64},"It applies NLP techniques in the first pipeline step, including sentence-level embeddings from models such as BERT and Word2Vec and multiple tokenization strategies to better handle the low-resource Text to Gloss setting.",{"name":70,"@type":61,"acceptedAnswer":71},"Why introduce Text to HamNoSys (T2H), and what extra supervision is used?",{"text":72,"@type":64},"The paper argues for using the phonetic HamNoSys representation instead of gloss-level representations. It also extracts hand shape from HamNoSys and uses it as additional supervision during training.","https://schema.org",{"og:url":32,"og:type":75,"og:title":10,"og:site_name":45,"og:description":12},"article",{"robots":77,"canonical":32},"index,follow",{"doc_id":79,"site_id":7},138300,1787478715,{"code":4,"msg":82,"data":83},"success",[84,88,92,96,101,106,111,115,119,122,126],{"id":22,"doc_module":4,"doc_module_name":25,"category_name":85,"show_sort_weight":86,"slug":87},"Story & Novel",90,"story-novel",{"id":26,"doc_module":4,"doc_module_name":25,"category_name":89,"show_sort_weight":90,"slug":91},"Literature",80,"literature",{"id":33,"doc_module":4,"doc_module_name":25,"category_name":93,"show_sort_weight":94,"slug":95},"Exam",70,"exam",{"id":97,"doc_module":4,"doc_module_name":25,"category_name":98,"show_sort_weight":99,"slug":100},5,"Comic",60,"comic",{"id":102,"doc_module":4,"doc_module_name":25,"category_name":103,"show_sort_weight":104,"slug":105},6,"Technology",50,"technology",{"id":107,"doc_module":4,"doc_module_name":25,"category_name":108,"show_sort_weight":109,"slug":110},7,"Healthcare",40,"healthcare",{"id":112,"doc_module":4,"doc_module_name":25,"category_name":29,"show_sort_weight":113,"slug":114},8,30,"research-report",{"id":55,"doc_module":4,"doc_module_name":25,"category_name":116,"show_sort_weight":117,"slug":118},"Religion & Spirituality",20,"religion-spirituality",{"id":117,"doc_module":4,"doc_module_name":25,"category_name":120,"show_sort_weight":117,"slug":121},"World Cup","world-cup",{"id":123,"doc_module":4,"doc_module_name":25,"category_name":124,"show_sort_weight":123,"slug":125},10,"Lifestyle","lifestyle",{"id":127,"doc_module":4,"doc_module_name":25,"category_name":128,"show_sort_weight":97,"slug":129},19,"General","general",{"code":4,"msg":82,"data":131},{"doc_id":79,"user_id":132,"nickname":42,"user_avatar":133,"doc_module":4,"category_id":112,"category_name":29,"doc_title":10,"doc_description":12,"doc_content":134,"file_id":135,"file_url":136,"file_type":137,"file_size":138,"view_count":55,"is_deleted":4,"is_public":22,"is_downloadable":22,"audit_status":22,"page_count":112,"language":139,"language_code":8,"site_id":7,"html_lang":8,"table_of_contents":140,"faqs":141,"seo_title":142,"seo_description":12,"update_tm":80,"read_time":117},7971461740909,"https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d","Proceedings of the 7th International Workshop on Sign Language Translation and Avatar Technology (SLTAT 7) , pages 117–124 Language Resources and Evaluation Conference (LREC 2022), Marseille, 20-25 June 2022  \n© European Language Resources Association (ELRA), licensed under CC-BY-NC 4.0  \nChanging the Representation: Examining Language Representation for  \nNeural Sign Language Production  \nHarry Walsh, Ben Saunders, Richard Bowden  \nUniversity of Surrey  \n{harry.walsh, b.saunders, [r.bowden](r.bowden}@surrey.ac.uk)[}](r.bowden}@surrey.ac.uk)[@surrey.ac.uk](r.bowden}@surrey.ac.uk)  \nAbstract  \nNeural Sign Language Production (SLP) aims to automatically translate from spoken language sentences to sign language videos. Historically the SLP task has been broken into two steps; Firstly, translating from a spoken language sentence to a gloss sequence and secondly, producing a sign language video given a sequence of glosses. In this paper we apply Natural Language Processing techniques to the first step of the SLP pipeline. We use language models such as BERT and Word2Vec to create better sentence level embeddings, and apply several tokenization techniques, demonstrating how these improve performance on the low resource translation task of Text to Gloss. We introduce Text to HamNoSys (T2H) translation, and show the advantages of using a phonetic representation for sign language translation rather than a sign level gloss representation. Furthermore, we use HamNoSys to extract the hand shape of a sign and use this as additional supervision during training, further increasing the performance on T2H. Assembling best practise, we achieve a BLEU-4 score of 26.99 on the MineDGS dataset and 25.09 on PHOENIX14T, two new state-of-the-art baselines.  \nKeywords: Sign Language Translation (SLT), Natural Language Processing (NLP), Sign Language, Phonetic Representation  \n1. Introduction  \nSign languages are the dominant form of communication for Deaf communities, with 430 million users worldwide (WHO, 2021) . Sign languages are complex multichannel languages with their own grammatical structure and vocabulary (Stokoe, 1980) . For many people, sign language is their primary language, and written forms of spoken language are their secondary languages.  \nSign Language Production (SLP) aims to bridge the gap between hearing and Deaf communities, by translating from spoken language sentences to sign language sequences. This problem has historically been broken into two steps; 1) translation from spoken language to gloss 1 and 2) subsequent production of sign language sequences from a sequence of glosses, commonly using a graphical avatar (Elliott et al., 2008; Efthimiou et al., 2010; Efthimiou et al., 2009) or more recently, a photorealistic signer (Saunders et al., 2021a; Saunders et al., 2021b) . In this paper, we improve the SLP pipeline by focusing on the Text to Gloss (T2G) translation task of step 1 .  \nModern deep learning is heavily dependent upon data. However, the creation of sign language datasets is both time consuming and costly, restricting their size to orders of magnitude smaller than their spoken language counterparts. State-of-the-art datasets such as RWTH-PHOENIX-Weather-2014T (PHOENIX14T), and the newer MineDGS (mDGS), contain only 8,257 and 63,912 examples respectively (Koller et al., 2015; Hanke et al., 2020), compared to over 15 million exam-  \n1 Gloss is the written word associated with a sign  \nples for common spoken language datasets (Vrandei and Krtzsch, 2014) . Hence, sign languages can be considered as low resource languages.  \nIn this work, we take inspiration from NLP techniques to boost translation performance. We explore how language can be modeled using different tokenizers, more specifically Byte Pair Encoding (BPE), WordPiece, word and character level tokenizers. We show that finding the correct tokenizer for the task helps simplify the translation problem.  \nFurthermore, to help tackle our low resource language task","cbCaidKJL78jLJiB","https://ap.wps.com/l/cbCaidKJL78jLJiB","pdf",646693,"English","# Introduction\n# Related Work","[{\"question\":\"What does Neural Sign Language Production (SLP) aim to do?\",\"answer\":\"SLP converts spoken-language sentences into sign language videos automatically. The task has traditionally been split into Text-to-Gloss translation and then video production from gloss sequences.\"},{\"question\":\"How does this paper improve the Text-to-Gloss step in SLP?\",\"answer\":\"It applies NLP techniques in the first pipeline step, including sentence-level embeddings from models such as BERT and Word2Vec and multiple tokenization strategies to better handle the low-resource Text to Gloss setting.\"},{\"question\":\"Why introduce Text to HamNoSys (T2H), and what extra supervision is used?\",\"answer\":\"The paper argues for using the phonetic HamNoSys representation instead of gloss-level representations. It also extracts hand shape from HamNoSys and uses it as additional supervision during training.\"}]","Changing the Representation - Examining Language Representation for Neural Sign Language Production | PDF"]