[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86174-en":3,"doc-seo-86174-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},86174,13056703019662,"Evangeline","https://ap-avatar.wpscdn.com/avatar/be000253a8e92610077?_k=1778726343310543188",8,"Research & Report","Q-BridgeNet: A Quantization Network for Cross-Lingual Sign Language Translation","Most sign language translation (SLT) research targets isolated sign–spoken pairs, leaving multilingual scenarios difficult to scale for real accessibility needs. Existing multilingual SLT methods often fail to learn a unified representation: they struggle to suppress cross-lingual conflicts while sharing semantics across languages and preserving language-specific variations. Q-BridgeNet addresses this by using adaptive segmentation and residual vector quantization to form discrete Q-units on the sign side, then fine-tuning a multilingual LLM to translate within that space, achieving state-of-the-art results and strong generalization.","arXiv :2607 . 1 12 15v 1 [ cs .CL] 13 Jul 2026  \nQ-BridgeNet: A Quantization Network for Cross-Lingual Sign Language Translation  \nLiqian Feng 1 , Lintao Wang 1 , Xiaochen Liu 1 , Anusha Withana 1 , Ken-Tye Yong 1 , Dehui Kong2 , Zhiyong Wang 1 , and Kun Hu3 ,✉  \n1 The University of Sydney, Darlington NSW, Australia  \n{lfen0902, [lwan3720}@uni.sydney.edu.au](lwan3720}@uni.sydney.edu.au)[ ](lwan3720}@uni.sydney.edu.au){xiaochen.liu, anusha.withana, ken.yong, [zhiyong.wang}@sydney.edu.au](zhiyong.wang}@sydney.edu.au)  \n2 Beijing University of Technology, Beijing, China  \n[kdh@bjut.edu.cn](kdh@bjut.edu.cn)  \n3 Edith Cowan University, Joondalup WA, Australia  \n[k.hu@ecu.edu.au](k.hu@ecu.edu.au)  \nAbstract. Most sign language translation (SLT) methods focus on isolated native sign–spoken pairs (e.g., American Sign Language–English) .  \nExtending language-specific SLT models to multilingual translation would improve accessibility by enabling communication across diverse sign and spoken language communities. However, existing multilingual SLT approaches still struggle to learn a unified model that minimizes crosslingual conflicts while capturing shared cross-lingual semantics and preserving language-specific variations across different sign languages. Therefore, we propose Q-BridgeNet, a unified framework for multilingual SLT that jointly mitigates cross-lingual conflicts across both the sign language and spoken language sides. On the sign language side, QBridgeNet learns discrete Q-units via adaptive segmentation and residual vector quantization: a shared base codebook provides language-agnostic semantic primitives, while language-specific residual codebooks refine heterogeneous signing semantics. On the spoken language side, a multilingual LLM is fine-tuned to operate in the Q-unit space, leveraging cross-lingual priors to enable a unified SLT model. Experiments on PHOENIX14T, How2Sign, and CSL-Daily show that Q-BridgeNet effectively mitigates cross-lingual conflicts, achieving state-of-the-art performance on native sign–spoken pairs while also demonstrating strong generalization to non-native pairs. Our source code is publicly available at: [https://github.com/FengLiQ/Q-BridgeNet](https://github.com/FengLiQ/Q-BridgeNet)  \nKeywords: Sign Language Translation · Representation Learning · Multilingual Modeling  \n1 Introduction  \nOver 360 million people worldwide rely on sign languages for daily communication [24] . As visual languages with their own phonology, syntax, and discourse  \n✉ Corresponding author.  \n2 L. Feng et al.  \nFig. 1: Multilingual SLT and signing units. (a) Monolingual SLT translates a single sign–spoken language pair. (b) Multilingual SLT extends to many-to-many translation across multiple sign and spoken languages. (c) Fixed-length segmentation may cut across semantic/lexical boundaries, mixing heterogeneous signing patterns and yielding poor alignment to text. (d) Variable-length semantic Q-Unit forms coherent signing units that better align with textual units and helps reduce cross-lingual interference.  \nstructure, sign languages differ fundamentally from spoken languages [23] . Bridging the two is therefore a core challenge in multimodal language understanding, with direct implications for accessibility and inclusion. Most prior work on Sign Language Translation (SLT) [1, 2, 8, 15, 20, 43] focuses on a single sign-spoken language pair, learning to map continuous sign observations (e.g., video/pose sequences) into spoken-language text. These approaches have achieved substantial progress in monolingual SLT by learning mappings between visual sign representations and textual outputs. While these methods have made strong progress in native (in-pair) settings, practical applications often require interaction across multiple sign languages (Figure 1 a,b)–e.g. , American Sign Language (ASL), Chinese Sign Language (CSL), German Sign Language (DGS)–and multiple spoken languages.  \nThis motivates multilingual SLT, which ","cbCaipGTgSp6VDnt","https://ap.wps.com/l/cbCaipGTgSp6VDnt","pdf",5991403,1,18,"English","en",105,"# Introduction\n## Multilingual SLT motivation\n## Cross-lingual conflicts in unified modeling\n# Q-BridgeNet framework\n## Sign-side discrete Q-unit learning\n## Spoken-side LLM fine-tuning\n# Experiments and datasets","[{\"question\":\"What problem does Q-BridgeNet address in multilingual sign language translation?\",\"answer\":\"It targets the difficulty of learning a unified multilingual SLT representation that reduces cross-lingual conflicts while preserving shared semantics and language-specific variations.\"},{\"question\":\"How does Q-BridgeNet represent signing units on the sign language side?\",\"answer\":\"It uses adaptive temporal segmentation to create variable-length semantic units, then applies residual vector quantization to generate discrete Q-units from shared base and language-specific residual codebooks.\"},{\"question\":\"How is the translation model built for spoken language outputs?\",\"answer\":\"A multilingual LLM is fine-tuned to operate directly on the resulting discrete Q-unit sequences, producing a unified many-to-many translation model across sign and spoken languages.\"}]",1784209111,45,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"q-bridgenet-a-quantization-network-for-cross-lingual-sign-language-translation","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/q-bridgenet-a-quantization-network-for-cross-lingual-sign-language-translation/86174/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does Q-BridgeNet address in multilingual sign language translation?","Question",{"text":75,"@type":76},"It targets the difficulty of learning a unified multilingual SLT representation that reduces cross-lingual conflicts while preserving shared semantics and language-specific variations.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does Q-BridgeNet represent signing units on the sign language side?",{"text":80,"@type":76},"It uses adaptive temporal segmentation to create variable-length semantic units, then applies residual vector quantization to generate discrete Q-units from shared base and language-specific residual codebooks.",{"name":82,"@type":73,"acceptedAnswer":83},"How is the translation model built for spoken language outputs?",{"text":84,"@type":76},"A multilingual LLM is fine-tuned to operate directly on the resulting discrete Q-unit sequences, producing a unified many-to-many translation model across sign and spoken languages.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]