[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86448-en":3,"doc-seo-86448-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86448,2336464648322,"Aria","https://ap-avatar.wpscdn.com/avatar/2200025388227c56fec?_k=1778556882303663488",8,"Research & Report","Do Transformer Temporal Heads and Post-Pooling Motion Gates Help CorrNet-based CSLR? An Empirical Study","CorrNet offers a strong baseline for continuous sign language recognition (CSLR) by capturing inter-frame correlations within the visual encoding stage. This study tests two natural architectural extensions on a reproduced CorrNet system: swapping the BiLSTM temporal head for a Transformer encoder and adding motion cues after temporal pooling via a lightweight MotionGate module. Experiments show the Transformer head fails to beat the BiLSTM baseline with comparable runtime cost, while MotionGate converges to an identity-like, weak residual with lost motion selectivity.","arXiv :2607 .09890v1 [ cs .CV] 10 Jul 2026  \nDo Transformer Temporal Heads and Post-Pooling Motion Gates Help CorrNet-based CSLR? An Empirical Study  \nLisi Wang 1 , Zhidong Xiao2 ,∗ , Jianjun Peng3  \n1University of New South Wales, Sydney, Australia  \n2National Centre for Computer Animation, Faculty of Media, Science and Technology,  \nBournemouth University, BH12 5BB, United Kingdom  \n3 School of Information Science and Engineering, Dalian Polytechnic University,  \nDalian 116034, China  \n[lisiwang.tech@gmail.com](lisiwang.tech@gmail.com)  \n[zxiao@bournemouth.ac.uk](zxiao@bournemouth.ac.uk), [pengjj@dlpu.edu.cn](pengjj@dlpu.edu.cn)  \n∗ Corresponding author  \nAbstract  \nCorrNet is a strong baseline for continuous sign language recognition (CSLR) because it models inter-frame correlations inside the visual encoding stage. In this paper, we study two natural extensions of a reproduced CorrNet system: replacing the BiLSTM temporal head with a Transformer encoder, and injecting motion cues after temporal pooling. We find that the Transformer head does not outperform the BiLSTM baseline, even with a training strategy adjusted for the Transformer, and the two heads have almost the same computational and runtime cost. For the second extension, we design a lightweight module called MotionGate. In our experiments, MotionGate consistently collapses to an identity-like mapping: the gate loses motion selectivity, and the injected residual becomes a weak, non-selective perturbation of the pooled features. These results suggest that explicit motion injection after CorrNet’s correlation-based encoding is largely redundant, and that natural-looking architectural extensions in CSLR should be tested carefully instead of being assumed to help.  \n1 Introduction  \nContinuous sign language recognition (CSLR) aims to recognize a sequence of glosses from an input signing video. Unlike isolated sign recognition, CSLR has to model both visual appearance and temporal dependencies over long video sequences, and training usually only uses sentence-level gloss annotations. Because of this, most modern CSLR systems follow a CTC-based pipeline: a visual backbone extracts frame-wise features, a temporal module models the resulting sequence, and CTC supervision aligns the predicted gloss sequence without frame-level labels.  \nA central challenge in CSLR is temporal modeling. Sign language contains rich motion patterns, such as hand trajectories, body movement, facial expressions, and transitions between signs. CorrNet handles this by modeling inter-frame correlations inside the visual encoder, which gives a strong representation for the downstream temporal module. This makes CorrNet a good basis for asking whether extra temporal or motion modules can still add useful information.  \nWe study two natural extensions on a reproduced CorrNet system. The first replaces the original BiLSTM temporal module with a lightweight Transformer encoder. Self-attention can model global dependencies directly, so it is often seen as a stronger temporal model than recurrent networks. However, in the CorrNet pipeline, the sequence after downsampling is already short, and the backbone has already encoded local motion correlations. It is therefore not clear whether a Transformer head can bring extra gain over the recurrent baseline.  \nThe second extension injects motion cues after temporal pooling. Early pooling may weaken fine-grained motion information, so post-pooling motion compensation looks like a reasonable fix. Based on this idea, we design a lightweight post-pooling motion gate, called MotionGate, which uses temporal feature differences to add residual motion information to the pooled sequence. We do not assume this module is helpful. Instead, we check whether the learned module actually contributes meaningful residual information in the reproduced CorrNet setting.  \nOur experiments show that neither extension improves recognition accuracy in the reproduced CorrNet setting, w","cbCaibxZMHfQwoD5","https://ap.wps.com/l/cbCaibxZMHfQwoD5","pdf",315902,2,1,12,"English","en",105,"# Introduction\n# Related Work\n## Continuous Sign Language Recognition\n## CorrNet and VAC","[{\"question\":\"What two extensions does the paper evaluate on a reproduced CorrNet-based CSLR system?\",\"answer\":\"The paper evaluates (1) replacing the BiLSTM temporal head with a Transformer encoder, and (2) injecting motion cues after temporal pooling using a lightweight MotionGate module.\"},{\"question\":\"Why does the paper consider post-pooling motion injection as a plausible improvement?\",\"answer\":\"It argues that early pooling may weaken fine-grained motion information, so adding motion compensation after pooling could restore useful motion cues.\"},{\"question\":\"What do the experiments conclude about recognition accuracy and MotionGate behavior?\",\"answer\":\"Neither extension improves recognition accuracy in the reproduced CorrNet setting, and MotionGate converges to an identity-like residual regime where the gate loses motion selectivity, becoming a weak, non-selective perturbation.\"}]",1784211799,30,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"do-transformer-temporal-heads-and-post-pooling-motion-gates-help-corrnet-based-cslr-an-empirical-study","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/do-transformer-temporal-heads-and-post-pooling-motion-gates-help-corrnet-based-cslr-an-empirical-study/86448/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-26","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What two extensions does the paper evaluate on a reproduced CorrNet-based CSLR system?","Question",{"text":75,"@type":76},"The paper evaluates (1) replacing the BiLSTM temporal head with a Transformer encoder, and (2) injecting motion cues after temporal pooling using a lightweight MotionGate module.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Why does the paper consider post-pooling motion injection as a plausible improvement?",{"text":80,"@type":76},"It argues that early pooling may weaken fine-grained motion information, so adding motion compensation after pooling could restore useful motion cues.",{"name":82,"@type":73,"acceptedAnswer":83},"What do the experiments conclude about recognition accuracy and MotionGate behavior?",{"text":84,"@type":76},"Neither extension improves recognition accuracy in the reproduced CorrNet setting, and MotionGate converges to an identity-like residual regime where the gate loses motion selectivity, becoming a weak, non-selective perturbation.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":29,"slug":121},"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]