[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82208-en":3,"doc-seo-82208-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},82208,1374391974468,"Eden","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","VTaMo Video Text Alignment Model for Sign Language Translation","Sign language translation (SLT) converts continuous sign videos into spoken-language text, where gloss-free systems often learn cross-modal alignment implicitly from translation supervision alone. VTaMo introduces explicit multi-granularity alignment at local, global, and token-position levels: local entropy-regularized optimal transport with a learnable null token, global embedding-geometry calibration via a learnable orthogonal transformation guided by Earth Mover’s Distance, and position-aligned contrastive learning. Experiments on Phoenix-2014T, CSL-Daily, How2Sign, and OpenASL show state-of-the-art results, with ablations validating each component’s complementarity.","arXiv :2607 .09126v1 [ cs .CV] 10 Jul 2026  \nVTaMo: Video-Text Alignment Model for Sign Language Translation  \nJunyi Hu 1 , Zhewen He 1 , Haomian Huang 1 , Aoxiang Yang 1 , and Yi Fang 1 ,2⋆  \n1 New York University Abu Dhabi, UAE  \n2 ChatSign Technology  \n[jh10472@nyu.edu](jh10472@nyu.edu), [yf23@nyu.edu](yf23@nyu.edu)  \nAbstract. Sign language translation (SLT) converts continuous sign videos into spoken language text. Gloss-free approaches leverage pretrained visual encoders and language models but rely on implicit crossmodal alignment from translation supervision alone. We present VTaMo, a framework that introduces explicit multi-granularity alignment at three levels: (1) local alignment via entropy-regularized optimal transport with a learnable null token for fine-grained frame-to-token correspondences;  \n(2) global alignment via a learnable orthogonal transformation that cali  \nbrates embedding space geometry through Earth Mover’s Distance; and  \n(3) position-aligned contrastive learning for discriminative token-level representations. Experiments on Phoenix-2014T, CSL-Daily, How2Sign, and OpenASL demonstrate consistent state-of-the-art performance, with ablations confirming the complementary contributions of each component. Code is available at [https://github.com/junyi2005/vtamo](https://github.com/junyi2005/vtamo).  \nKeywords: Sign language translation · Cross-modal alignment · Optimal transport  \n1 Introduction  \nSign language translation (SLT) aims to convert continuous sign language videos into spoken language sentences, enabling more accessible communication between Deaf and hard-of-hearing signers and hearing communities [1, 48] . Recent gloss-free systems decode text directly from visual features using large pretrained sequence-to-sequence language models [7,13,40], avoiding the costly gloss annotation required by gloss-based pipelines. However, most gloss-free benchmarks provide only natural language sentences as supervision, without gloss annotations or temporal segmentation [11, 36] . The core difficulty is therefore twofold: the semantic gap between vision and text, and the unknown correspondence between the temporal order of visual signs and the token order of the target text.  \nSign languages often do not follow the word order of the corresponding spoken language: a signer may express key content words first and add grammatical  \n⋆ Corresponding author.  \n2 J. Hu et al.  \nFig. 1: Motivation and performance of VTaMo. (a) Prior gloss-free sign language translation methods typically rely on implicit cross-modal alignment learned inside the decoder attention, which can yield diffuse or mismatched cross-attention and translation errors. (b) VTaMo introduces explicit multi-granularity vision–text alignment—including local OT-based token-to-frame matching, global distribution alignment, and position-aligned contrastive learning—to sharpen correspondences and improve decoding. (c) VTaMo achieves state-of-the-art sign language translation quality across four benchmarks (Phoenix-2014T, CSL-Daily, How2Sign, and OpenASL), compared against SpaMo [16], Uni-Sign [21], SHuBERT [14], and SSVP-SLT [33] . Video frames in (a) and (b) are from the How2Sign dataset [11] .  \nrelations later, so the temporal progression of gestures can differ from the target word order. As a result, visual features extracted along the video timeline are often misaligned with the text tokens that an autoregressive decoder is trained to predict. Left unresolved, the decoder must simultaneously learn translation and implicitly discover a latent cross-modal permutation, which increases optimization difficulty and degrades accuracy, especially on large-scale benchmarks such as How2Sign and OpenASL. Fig. 1 illustrates this gap: when alignment is left implicit, the cross-attention map can be noisy and the decoder may produce incorrect word ordering, whereas our approach makes alignment explicit.  \nExisting methods largely treat the visual stream as an ord","cbCaidRXud79HjNo","https://ap.wps.com/l/cbCaidRXud79HjNo","pdf",2056789,1,27,"English","en",105,"# Introduction\n## Problem of implicit alignment in gloss-free SLT\n## Related work and limitations\n## Proposed VTaMo approach","[{\"question\":\"What does VTaMo improve compared with gloss-free sign language translation methods?\",\"answer\":\"VTaMo makes cross-modal alignment explicit instead of relying mainly on decoder attention, using multi-granularity alignment to sharpen correspondences and improve decoding quality.\"},{\"question\":\"How does VTaMo perform local alignment between video frames and target tokens?\",\"answer\":\"It estimates a soft matching with entropy-regularized optimal transport and a learnable null token to handle transitional gestures that do not correspond to explicit words.\"},{\"question\":\"Which benchmarks were used to evaluate VTaMo, and what overall outcome was reported?\",\"answer\":\"VTaMo was evaluated on Phoenix-2014T, CSL-Daily, How2Sign, and OpenASL, reporting consistent state-of-the-art performance, with ablations confirming contributions from each alignment component.\"}]",1784178820,68,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"vtamo-video-text-alignment-model-for-sign-language-translation","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/vtamo-video-text-alignment-model-for-sign-language-translation/82208/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What does VTaMo improve compared with gloss-free sign language translation methods?","Question",{"text":75,"@type":76},"VTaMo makes cross-modal alignment explicit instead of relying mainly on decoder attention, using multi-granularity alignment to sharpen correspondences and improve decoding quality.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does VTaMo perform local alignment between video frames and target tokens?",{"text":80,"@type":76},"It estimates a soft matching with entropy-regularized optimal transport and a learnable null token to handle transitional gestures that do not correspond to explicit words.",{"name":82,"@type":73,"acceptedAnswer":83},"Which benchmarks were used to evaluate VTaMo, and what overall outcome was reported?",{"text":84,"@type":76},"VTaMo was evaluated on Phoenix-2014T, CSL-Daily, How2Sign, and OpenASL, reporting consistent state-of-the-art performance, with ablations confirming contributions from each alignment component.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]