[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-134751-en":3,"doc-seo-134751-105":31,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":30},134751,1374391975076,"Riley","https://ap-avatar.wpscdn.com/avatar/14000253ca4ec9f6853?x-image-process=image/resize,m_fixed,w_180,h_180&k=1783305029341752051",8,"Research & Report","CVT-SLR - Contrastive Visual-Textual Transformation for Sign Language Recognition","Sign language recognition (SLR) is a weakly supervised multimodal task that annotates sign videos with textual glosses, yet progress is constrained by the scarcity of large-scale sign-language datasets such as PHOENIX-2014 and PHOENIX-2014T. This paper addresses the bottleneck by leveraging pretrained knowledge from both visual and language modalities through a single-cue cross-modal alignment design. It introduces CVT-SLR with a variational autoencoder to model pretrained contextual knowledge and a contrastive alignment algorithm to strengthen cross-modal consistency. Experiments on public datasets show consistent improvements over existing single-cue methods and even outperform multi-cue state-of-the-art approaches.","This CVPR paper is the Open Access version, provided by the Computer Vision Foundation.  \nExcept for this watermark, it is identical to the accepted version; the final published version of the proceedings is available on IEEE Xplore.  \nCVT-SLR: Contrastive Visual-Textual Transformation for Sign Language Recognition with Variational Alignment  \nJiangbin Zheng1 , Yile Wang1,2 , Cheng Tan1 , Siyuan Li1 , Ge Wang1 , Jun Xia1 , Yidong Chen3 , Stan Z. Li1 *  \n1AI Lab, Research Center for Industries of the Future, Westlake University  \n2Institute for AI Industry Research (AIR), Tsinghua University  \n3 School of Informatics, Xiamen University  \n{zhengjiangbin,wangyile,tancheng,lisiyuan,wangge,xiajun,[Stan.ZQ.Li](Stan.ZQ.Li}@westlake.edu.cn)[}](Stan.ZQ.Li}@westlake.edu.cn)[@westlake.edu.cn](Stan.ZQ.Li}@westlake.edu.cn)  \n[ydchen@xmu.edu.cn](ydchen@xmu.edu.cn)  \nAbstract  \nSign language recognition (SLR) is a weakly supervised task that annotates sign videos as textual glosses. Recent studies show that insufficient training caused by the lack of large-scale available sign datasets becomes the main bottleneck for SLR. Most SLR works thereby adopt pretrained visual modules and develop two mainstream solutions. The multi-stream architectures extend multi-cue visual features, yielding the current SOTA performances but requiring complex designs and might introduce potential noise. Alternatively, the advanced single-cue SLR frameworks using explicit cross-modal alignment between visual and textual modalities are simple and effective, potentially competitive with the multi-cue framework. In this work, we propose a novel contrastive visual-textual transformation for SLR, CVT-SLR, to fully explore the pretrained knowledge of both the visual and language modalities. Based on the single-cue cross-modal alignment framework, we propose a variational autoencoder (VAE) for pretrained contextual knowledge while introducing the complete pretrained language module. The VAE implicitly aligns visual and textual modalities while benefiting from pretrained contextual knowledge as the traditional contextual module. Meanwhile, a contrastive cross-modal alignment algorithm is designed to explicitly enhance the consistency constraints. Extensive experiments on public datasets (PHOENIX-2014 and PHOENIX-2014T) demonstrate that our proposed CVT-SLR consistently outperforms existing single-cue methods and even outperforms SOTA multi-cue methods. The source codes and models are available at [https://github. com/binbinjiang/CVT-SLR](https://github. com/binbinjiang/CVT-SLR).  \n*Corresponding author.  \n1. Introduction  \nAs a special visual natural language, sign language is the primary communication medium of the deaf community [19] . With the progress of deep learning [1, 17, 25, 39, 42], sign language recognition (SLR) has emerged as a multimodal task that aims to annotate sign videos into textual sign glosses. However, a significant dilemma of SLR is the lack of publicly available sign language datasets. For example, the most commonly-used PHOENIX-2014 [23] and PHOENIX-2014T [2] datasets only include about 10K pairs of sign videos and gloss annotations, which are far from training a robust SLR system with full supervision as typical vision-language cross-modal tasks [34] . Therefore, data limitation that may easily lead to insufficient training or overfitting problems is the main bottleneck of SLR tasks.  \nThe development of weakly supervised SLR has witnessed most of the improvement efforts focus on the visual module (e.g., CNN) [9, 10, 15, 29, 32, 33] . Transferring pretrained visual networks from general domains of human actions becomes a consensus to alleviate the low-resource limitation. The mainstream multi-stream SLR framework extends the pretrained visual module with multi-cue visual information [3, 22, 24,43,48, 50], including global features and regional features such as hands and faces in independent streams. The theoretical support for this approach comes","cbCair0PYw8vx5Ph","https://ap.wps.com/l/cbCair0PYw8vx5Ph","pdf",870803,2,1,10,"English","en",105,"# Abstract\n# 1. Introduction\n## Sign language recognition and dataset limitations\n## Multi-stream SLR and its challenges\n## Single-cue SLR with explicit cross-modal alignment","[{\"question\":\"What problem does CVT-SLR target in sign language recognition?\",\"answer\":\"CVT-SLR targets the main bottleneck of SLR: insufficient training caused by the lack of large-scale publicly available sign-language datasets.\"},{\"question\":\"How does CVT-SLR differ from typical single-cue cross-modal alignment frameworks?\",\"answer\":\"It replaces the mainstream contextual module with a variational autoencoder (VAE) to leverage pretrained contextual knowledge while enabling implicit cross-modal alignment.\"},{\"question\":\"What evidence shows the effectiveness of CVT-SLR?\",\"answer\":\"Extensive experiments on PHOENIX-2014 and PHOENIX-2014T demonstrate that CVT-SLR consistently outperforms existing single-cue methods and can surpass state-of-the-art multi-cue methods.\"}]","CVT-SLR - Contrastive Visual-Textual Transformation for Sign Language Recognition | PDF",1787299024,25,{"code":4,"msg":32,"data":33},"ok",{"site_id":25,"language":24,"slug":34,"title":13,"keywords":35,"description":14,"schema_data":36,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":29},"cvt-slr-contrastive-visual-textual-transformation-for-sign-language-recognition","",{"@graph":37,"@context":86},[38,54,69],{"@type":39,"itemListElement":40},"BreadcrumbList",[41,45,48,51],{"item":42,"name":43,"@type":44,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":46,"name":47,"@type":44,"position":20},"https://docshare.wps.com/document/","Document",{"item":49,"name":12,"@type":44,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":44,"position":53},"https://docshare.wps.com/document/cvt-slr-contrastive-visual-textual-transformation-for-sign-language-recognition/134751/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":42,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-23","2026-08-21",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What problem does CVT-SLR target in sign language recognition?","Question",{"text":76,"@type":77},"CVT-SLR targets the main bottleneck of SLR: insufficient training caused by the lack of large-scale publicly available sign-language datasets.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does CVT-SLR differ from typical single-cue cross-modal alignment frameworks?",{"text":81,"@type":77},"It replaces the mainstream contextual module with a variational autoencoder (VAE) to leverage pretrained contextual knowledge while enabling implicit cross-modal alignment.",{"name":83,"@type":74,"acceptedAnswer":84},"What evidence shows the effectiveness of CVT-SLR?",{"text":85,"@type":77},"Extensive experiments on PHOENIX-2014 and PHOENIX-2014T demonstrate that CVT-SLR consistently outperforms existing single-cue methods and can surpass state-of-the-art multi-cue methods.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,135],{"id":21,"doc_module":4,"doc_module_name":47,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":47,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":47,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":47,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":47,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":47,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":47,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":47,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":47,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":22,"doc_module":4,"doc_module_name":47,"category_name":133,"show_sort_weight":22,"slug":134},"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":47,"category_name":137,"show_sort_weight":107,"slug":138},19,"General","general"]