[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82899-en":3,"doc-seo-82899-105":28,"detail-sidebar-cat-0-en-105":90},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":11,"language":21,"language_code":22,"site_id":23,"html_lang":22,"table_of_contents":24,"faqs":25,"seo_title":13,"seo_description":14,"update_tm":26,"read_time":27},82899,8796095461564,"Liam","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","Video-Text Temporal Localization via Multi-Scale Convolution and Dynamic Routing","Video–text temporal localization aligns natural language queries with the precise temporal boundaries of video segments, a core problem in multimodal understanding. The framework addresses two key gaps in prior work: weak hierarchical modeling of temporal structure and limited capability for complex many-to-many modality correspondences. It combines a multi-scale temporal convolutional encoder for motion patterns at multiple granularities with capsule-based dynamic routing that iteratively refines segment–query associations via structured agreement updates. Joint multi-task learning optimizes boundary regression, cross-modal alignment, and capsule diversity. Experiments on ActivityNet Captions reach 42.9% Recall@0.5 and 41.1% mean IoU, outperforming strong transformer baselines while remaining computationally efficient.","Video-Text Temporal Localization via Multi-Scale Convolution and Dynamic  \nRouting  \nGengtian Shi 1 , Jinze Yu2,1 , Chenhao Wu 1 , Shaofei Wang 1 , Eiji Fukuzawa 1 , Junjie Tang3 , Hiroshi Onoda 1 , Jiang Liu 1  \n1 Graduate School of Fundamental Science and Engineering, Waseda University, Tokyo, Japan  \n2 Generative AI Innovation Center, Amazon Web Services, Japan  \n3Amazon, Germany  \n[shigengtian@akane.waseda.jp](shigengtian@akane.waseda.jp), [jinzeyu@amazon.co.jp](jinzeyu@amazon.co.jp)  \narXiv :2607 .05093v2 [ cs .CV] 7 Jul 2026  \nAbstract  \nVideo–text temporal localization requires precise alignment between natural language queries and corresponding video segments, a fundamental challenge in multimodal understanding. We present a novel framework that addresses two critical limitations of existing methods: inadequate modeling of hierarchical temporal structure and inability to handle complex many-to-many correspondences between modalities. Our approach introduces a multi-scale temporal convolutional encoder that captures motion patterns across different temporal granularities—from instantaneous frame transitions to extended action sequences. We further propose a capsule-based dynamic routing mechanism that iteratively refines segment–query associations through structured agreement updates, enabling flexible modeling of non-monotonic alignments. These components are unified through a multi-task learning objective that jointly optimizes temporal boundary regression, cross-modal semantic alignment, and capsule diversity. Extensive experiments on ActivityNet Captions demonstrate significant improvements, achieving 42.9% Recall@0.5 and 41.1% mean IoU, surpassing strong transformer-based baselines while maintaining computational efficiency. Our results validate that combining hierarchical temporal modeling with structured semantic routing provides an effective solution for fine-grained video–language understanding.  \nIntroduction  \nThe proliferation of video content has created an urgent need for systems that can understand and navigate temporal relationships between visual content and natural language. Video–text temporal localization, also known as temporal sentence grounding (TSG) or video moment retrieval (VMR), addresses this challenge by identifying precise temporal boundaries in videos that correspond to textual descriptions (Gao et al. 2017; Hendricks et al. 2017; Lan et al. 2023; Zhang et al. 2023) . This capability is essential for applications ranging from instructional video understanding and human–robot interaction to content-based video retrieval and interactive video editing (Chen et al. 2018; Liet al. 2020) .  \nDespite recent advances in vision–language modeling, achieving fine-grained temporal alignment remains fundamentally challenging. Current pretrained models such as CLIP (Radford et al. 2021), BLIP (Li et al. 2022), and their  \nvideo-adapted variants (e.g., CLIP4Clip (Luo et al. 2021), VideoBERT (Sun et al. 2019), X-CLIP (Ma et al. 2022)) excel at coarse-grained video–text matching but struggle with precise moment-level localization. These models typically operate on entire video clips or large temporal segments, lacking the temporal resolution necessary to identify exact action boundaries, subtle transitions, and complex visuallinguistic correspondences (Li et al. 2023) .  \nThe challenge is particularly acute in instructional and procedural videos, where multiple related actions occur inclose temporal proximity with ambiguous boundaries. Consider a cooking video where ”adding ingredients” and ”mixing the batter” may overlap temporally, or an assembly tutorial where distinct textual instructions correspond to partially concurrent visual actions. As illustrated in Figure 1, real-world scenarios frequently exhibit three types of complexity: (1) single queries spanning multiple disjoint video segments,(2) multiple queries mapping to overlapping temporal regions, and (3) subtle transitions between semantically re","cbCaibhC5JOwX5pH","https://ap.wps.com/l/cbCaibhC5JOwX5pH","pdf",1452429,1,"English","en",105,"# Introduction\n## Problem: Video–Text Temporal Localization\n## Limitations of Existing Methods\n## Proposed Framework and Key Ideas","[{\"question\":\"What problem does the paper address?\",\"answer\":\"The paper addresses video–text temporal localization (temporal sentence grounding), which finds precise video time boundaries that correspond to natural language descriptions.\"},{\"question\":\"What are the two main limitations of existing approaches highlighted in the paper?\",\"answer\":\"Most methods do not explicitly model hierarchical temporal structure across different time scales, and their alignment mechanisms assume overly simple (often monotonic or one-to-one) correspondences, failing on many-to-many relationships.\"},{\"question\":\"How does the proposed method improve temporal localization?\",\"answer\":\"It uses a multi-scale temporal convolutional encoder to capture motion at multiple granularities and capsule-based dynamic routing to iteratively refine segment–query associations through structured agreement updates, optimized jointly with multi-task learning objectives.\"}]",1784183802,20,{"code":4,"msg":29,"data":30},"ok",{"site_id":23,"language":22,"slug":31,"title":13,"keywords":32,"description":14,"schema_data":33,"social_meta":85,"head_meta":87,"extra_data":89,"updated_unix":26},"video-text-temporal-localization-via-multi-scale-convolution-and-dynamic-routing","",{"@graph":34,"@context":84},[35,52,67],{"@type":36,"itemListElement":37},"BreadcrumbList",[38,42,46,49],{"item":39,"name":40,"@type":41,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":43,"name":44,"@type":41,"position":45},"https://docshare.wps.com/document/","Document",2,{"item":47,"name":12,"@type":41,"position":48},"https://docshare.wps.com/document/research-report/",3,{"item":50,"name":13,"@type":41,"position":51},"https://docshare.wps.com/document/video-text-temporal-localization-via-multi-scale-convolution-and-dynamic-routing/82899/",4,{"url":50,"name":13,"@type":53,"author":54,"headline":13,"publisher":56,"fileFormat":59,"inLanguage":22,"description":14,"dateModified":60,"datePublished":61,"encodingFormat":59,"isAccessibleForFree":62,"interactionStatistic":63},"DigitalDocument",{"name":9,"@type":55},"Person",{"url":39,"name":57,"@type":58},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":64,"interactionType":65,"userInteractionCount":20},"InteractionCounter",{"@type":66},"ViewAction",{"@type":68,"mainEntity":69},"FAQPage",[70,76,80],{"name":71,"@type":72,"acceptedAnswer":73},"What problem does the paper address?","Question",{"text":74,"@type":75},"The paper addresses video–text temporal localization (temporal sentence grounding), which finds precise video time boundaries that correspond to natural language descriptions.","Answer",{"name":77,"@type":72,"acceptedAnswer":78},"What are the two main limitations of existing approaches highlighted in the paper?",{"text":79,"@type":75},"Most methods do not explicitly model hierarchical temporal structure across different time scales, and their alignment mechanisms assume overly simple (often monotonic or one-to-one) correspondences, failing on many-to-many relationships.",{"name":81,"@type":72,"acceptedAnswer":82},"How does the proposed method improve temporal localization?",{"text":83,"@type":75},"It uses a multi-scale temporal convolutional encoder to capture motion at multiple granularities and capsule-based dynamic routing to iteratively refine segment–query associations through structured agreement updates, optimized jointly with multi-task learning objectives.","https://schema.org",{"og:url":50,"og:type":86,"og:title":13,"og:site_name":57,"og:description":14},"article",{"robots":88,"canonical":50},"index,follow",{"doc_id":7,"site_id":23},{"code":4,"msg":5,"data":91},[92,96,100,104,109,114,119,122,126,129,133],{"id":20,"doc_module":4,"doc_module_name":44,"category_name":93,"show_sort_weight":94,"slug":95},"Story & Novel",90,"story-novel",{"id":45,"doc_module":4,"doc_module_name":44,"category_name":97,"show_sort_weight":98,"slug":99},"Literature",80,"literature",{"id":51,"doc_module":4,"doc_module_name":44,"category_name":101,"show_sort_weight":102,"slug":103},"Exam",70,"exam",{"id":105,"doc_module":4,"doc_module_name":44,"category_name":106,"show_sort_weight":107,"slug":108},5,"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":44,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":44,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":44,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":44,"category_name":124,"show_sort_weight":27,"slug":125},9,"Religion & Spirituality","religion-spirituality",{"id":27,"doc_module":4,"doc_module_name":44,"category_name":127,"show_sort_weight":27,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":44,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":44,"category_name":135,"show_sort_weight":105,"slug":136},19,"General","general"]