[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-140079-en":3,"doc-seo-140079-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},140079,13056703019404,"Miles","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",6,"Technology","Number-Prompt (NumPro) - Temporal Grounding Videos like Flipping Manga","Video Large Language Models (Vid-LLMs) excel at answering questions about video content, yet they often fail to provide precise temporal localization required by Video Temporal Grounding (VTG). To improve temporal boundary reasoning, the paper proposes Number-Prompt (NumPro), which overlays unique numerical identifiers onto each frame and treats the video as an ordered sequence of numbered images. This enables models to map visual evidence to timestamps in a manga-flipping style. Experiments show significant VTG gains without extra computation, and NumPro-enhanced fine-tuning achieves new state-of-the-art results.","This CVPR paper is the Open Access version, provided by the Computer Vision Foundation.  \nExcept for this watermark, it is identical to the accepted version; the final published version of the proceedings is available on IEEE Xplore.  \nNumber it: Temporal Grounding Videos like Flipping Manga  \nYongliang Wu 1 ,2 ,4 *† Wenbo Zhu5  \nXinting Hu3 * Fengyun Rao4  \nYuyang Sun 1 ,2 Yizhou Zhou4‡ Bernt Schiele3 Xu Yang 1 ,2§  \n1 Southeast University  \n2 Key Laboratory of New Generation Artiﬁcial Intelligence Technology and Its Interdisciplinary Applications (Southeast University), Ministry of Education, China  \n3Max Planck Institute for Informatics, Saarland Informatics Campus, Germany  \n4WeChat Vision, Tencent Inc. 5University of California, Berkeley  \n[yongliang0223@gmail.com](yongliang0223@gmail.com) xuyang [palm@seu.edu.cn](palm@seu.edu.cn)  \nFigure 1 . Effectiveness of Adding Frame Numbers for Temporal Grounding: (a) Without numbered images or frames, both humans and Vid-LLMs struggle to locate speciﬁc timestamps accurately. (b) Once numbered, grounding temporal cues becomes as intuitive as ﬂipping manga, where timestamps are accessible at a glance.  \nAbstract  \nVideo Large Language Models (Vid-LLMs) have made remarkable advancements in comprehending video content for QA dialogue. However, they struggle to extend this visual understanding to tasks requiring precise temporal localization, known as Video Temporal Grounding (VTG) . To address this, we introduce Number-Prompt (NumPro), a novel method that empowers Vid-LLMs to bridge visual comprehension with temporal grounding by adding unique numerical identiﬁers to each video frame. Treating a video as a sequence of numbered frame images, NumPro transforms VTG into an intuitive process: ﬂipping through manga panels in sequence. This allows Vid-LLMs to “read”event timelines, accurately linking visual content with cor-  \n* Equal Contribution. ‡ Project Leader. § Corresponding Author.† Work done during an internship at WeChat Vision, Tencent Inc.  \nresponding temporal information. Our experiments demonstrate that NumPro signiﬁcantly boosts VTG performance of top-tier Vid-LLMs without additional computational cost. Furthermore, ﬁne-tuning on a NumPro-enhanced dataset deﬁnes a new state-of-the-art for VTG, surpassing previous top-performing methods by up to 6.9% in mIoU for moment retrieval and 8. 5% in mAP for highlight detection. The code is available at [https://github.com/yongliang](https://github.com/yongliang)wu/NumPro.  \n1. Introduction  \nImagine you are watching a cooking video, and trying to locate the exact moment when the chef stirs in the spices. While recognizing such actions is feasible, translating that visual information into precise timing, i.e., a speciﬁc second or frame number, is surprisingly difﬁcult. This chal-  \nlenge is central to the ﬁeld of Video Temporal Grounding (VTG) [4, 18, 25, 36, 52, 58] . In the realm of Video Large Language Models (Vid-LLMs) [35, 43, 54, 66, 84, 89] which process videos as a sequence of frame images, the integration of VTG allows for ﬁne-grained visual and temporal understanding and reasoning of videos, which is pivotal for developing end-to-end video dialogue systems.  \nDespite advances of Vid-LLMs, endowing these models with effective VTG abilities presents a unique challenge: enhancing the model’s visual recognition of an event within a video does not inherently enable it to describe when the event begins and ends using language [25, 58] . For instance, advanced Vid-LLMs like Qwen2-VL [66], while excelling at video comprehension, can struggle with grounding speciﬁc events in time. When asked, e.g., to locate “when does the woman eat food” in a 10-frame video, the model can hallucinate an illogical answer like “from frame 000 to 580.”* This limitation arises because these models are primarily trained to align visual content with language descriptions (what happens) while lacking mechanisms to directly interpret the temporal boundaries","cbCaivO0OLSva2Cu","https://ap.wps.com/l/cbCaivO0OLSva2Cu","pdf",4005971,1,12,"English","en",105,"# Abstract\n# Introduction\n## Motivation and challenge of VTG for Vid-LLMs\n## Manga-inspired solution with frame numbers\n## NumPro design and temporal cue positioning\n## Key idea: no extra tokens or vocabulary changes","[{\"question\":\"What problem does NumPro address in Vid-LLMs?\",\"answer\":\"NumPro targets the difficulty Vid-LLMs have in converting visual understanding into precise temporal boundaries for Video Temporal Grounding (VTG).\"},{\"question\":\"How does Number-Prompt (NumPro) work?\",\"answer\":\"NumPro adds unique numerical identifiers to each video frame, turning the video into an ordered sequence of numbered images so models can associate frame cues with timestamps and output temporal information.\"},{\"question\":\"What benefits does NumPro bring according to the experiments?\",\"answer\":\"NumPro significantly improves VTG performance on top-tier Vid-LLMs without additional computational cost, and fine-tuning with a NumPro-enhanced dataset achieves a new state-of-the-art.\"}]","Number-Prompt (NumPro) - Temporal Grounding Videos like Flipping Manga | PDF",1787567517,30,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"number-prompt-numpro-temporal-grounding-videos-like-flipping-manga","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/technology/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/number-prompt-numpro-temporal-grounding-videos-like-flipping-manga/140079/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-24",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does NumPro address in Vid-LLMs?","Question",{"text":75,"@type":76},"NumPro targets the difficulty Vid-LLMs have in converting visual understanding into precise temporal boundaries for Video Temporal Grounding (VTG).","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does Number-Prompt (NumPro) work?",{"text":80,"@type":76},"NumPro adds unique numerical identifiers to each video frame, turning the video into an ordered sequence of numbered images so models can associate frame cues with timestamps and output temporal information.",{"name":82,"@type":73,"acceptedAnswer":83},"What benefits does NumPro bring according to the experiments?",{"text":84,"@type":76},"NumPro significantly improves VTG performance on top-tier Vid-LLMs without additional computational cost, and fine-tuning with a NumPro-enhanced dataset achieves a new state-of-the-art.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,113,118,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":111,"slug":112},50,"technology",{"id":114,"doc_module":4,"doc_module_name":46,"category_name":115,"show_sort_weight":116,"slug":117},7,"Healthcare",40,"healthcare",{"id":119,"doc_module":4,"doc_module_name":46,"category_name":120,"show_sort_weight":29,"slug":121},8,"Research & Report","research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]