[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82900-en":3,"doc-seo-82900-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},82900,8796095461564,"Liam","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","TimeThink: Reasoning with Time for Video LLMs","Video reasoning requires models to locate and validate temporally localized evidence within long video sequences. Recent Video-LLMs show strong reasoning under reinforcement learning, but prior methods often use outcome-only rewards that supervise only the final answer, giving weak intermediate guidance. TimeThink introduces a reinforcement learning framework that treats temporal clue steps as the core optimization unit, adding step-wise temporal process rewards and a joint process–outcome objective. TimeThink-RFT-20K supports scalable training via automatically derived temporal evidence segments, improving temporal grounding and reasoning performance.","arXiv :2607 .05089v 1 [ cs .CV] 6 Jul 2026  \nTimeThink: Reasoning with Time for Video LLMs  \nHandong Li 1 ,2⋆, Longteng Guo2⋆, Zikang Liu 1 ,2⋆, Dongze Hao3 , Yepeng Tang2 , Zijia Zhao2 , Jie Jiang2 , Zhiwei Jin3 , Chen Chen3 , Haonan Lu3 , and Jing  \nLiu 1 ,2⋆⋆  \n1 School of Artificial Intelligence, University of Chinese Academy of Sciences  \n2 Institute of Automation, Chinese Academy of Sciences  \n3 OPPO AI Center, OPPO Inc.  \n{[lihandong2023}@ia.ac.cn](lihandong2023}@ia.ac.cn) , {longteng.guo,[jliu}@nlpr.ia.ac.cn](jliu}@nlpr.ia.ac.cn)  \nAbstract. Video reasoning requires models to identify and verify temporally localized evidence within long video sequences. Recent Video Large Language Models (Video-LLMs) have shown promising reasoning abilities when aligned with reinforcement learning, yet existing approaches typically rely on outcome-based rewards that supervise only the final prediction. Such supervision provides limited guidance on how models should discover the relevant temporal evidence during intermediate reasoning. In this work, we propose TimeThink, a reinforcement learning framework that explicitly guides temporal evidence discovery in Video-LLMs. Our key idea is to treat temporal clue steps as the fundamental optimization primitive of video reasoning, where each reasoning step references a candidate time interval in the video. We introduce a step-wise temporal process reward that provides localized credit assignment for these clues and a joint process–outcome optimization objective that balances reasoning fidelity with task correctness. To enable scalable training, we construct TimeThink-RFT-20K, a dataset with automatically derived temporal evidence segments. Extensive experiments across video reasoning, temporal grounding, and general video understanding benchmarks show that TimeThink consistently improves both temporal localization and reasoning performance, achieving state-of-the-art results among open-source video RL models.  \nKeywords: Video-LLMs · Reinforcement Learning · Process Reward  \n1 Introduction  \nRecent advances in Video Large Language Models (Video-LLMs) [4, 28, 47, 67] have significantly expanded the frontier of machine video understanding. By integrating pretrained vision encoders with large language models, these systems are capable of performing diverse tasks such as video question answering, event narration, and complex video reasoning [9,30,52,71,72] . However, enabling models to reason reliably over long temporal contexts remains a central challenge.  \n⋆ Equal contribution.  \n⋆⋆ Corresponding author.  \n2 H. Li et al.  \nFig. 1: TimeThink vs. Traditional Video-CoT. (Left) Traditional Video-CoT relies on outcome-based rewards that supervise only the final answer, providing limited guidance for temporal evidence discovery. TimeThink instead models reasoning as a sequence of temporal clue steps, where each step references a candidate time interval in the video. (Right) Step-wise temporal process rewards supervise these clues, enabling (1) temporally grounded reasoning, (2) faster RL convergence, and (3) stronger video understanding.  \nUnlike static images, videos encode dynamic processes where events unfold overtime and critical evidence may appear only within brief segments of a long sequence. Consequently, effective video reasoning requires models to progressively inspect and verify temporally localized evidence rather than relying solely on global semantic impressions of the video.  \nTo improve reasoning capabilities, recent work has begun to apply reinforcement learning (RL) to Video-LLMs [15, 31, 49, 68] . Algorithms such as Group Relative Policy Optimization (GRPO) [44] enable models to explore multiple reasoning trajectories and optimize them using reward signals. In these frameworks, the model generates intermediate reasoning steps before producing a final answer, and the learning objective encourages trajectories that lead to correct predictions. While such approaches have shown promis","cbCaiqkrlAbJkVBV","https://ap.wps.com/l/cbCaiqkrlAbJkVBV","pdf",6542209,1,26,"English","en",105,"# Introduction\n## Motivation: temporal evidence in long videos\n## Limitation of existing RL with outcome-only rewards\n## TimeThink framework: temporal clue steps\n## Step-wise temporal process reward and joint optimization","[{\"question\":\"What problem does TimeThink target in video reasoning?\",\"answer\":\"TimeThink addresses the difficulty of guiding models to discover and verify temporally localized evidence across long video sequences during intermediate reasoning.\"},{\"question\":\"How does TimeThink differ from traditional outcome-based Video-CoT reinforcement learning?\",\"answer\":\"TimeThink models reasoning as a sequence of temporal clue steps, each referencing a candidate time interval, instead of relying mainly on rewards for only the final prediction.\"},{\"question\":\"What training signals does TimeThink use to improve learning?\",\"answer\":\"It introduces a step-wise temporal process reward for localized credit assignment and a joint process–outcome optimization objective to balance reasoning fidelity with task correctness.\"}]",1784183810,66,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"timethink-reasoning-with-time-for-video-llms","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/timethink-reasoning-with-time-for-video-llms/82900/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does TimeThink target in video reasoning?","Question",{"text":75,"@type":76},"TimeThink addresses the difficulty of guiding models to discover and verify temporally localized evidence across long video sequences during intermediate reasoning.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does TimeThink differ from traditional outcome-based Video-CoT reinforcement learning?",{"text":80,"@type":76},"TimeThink models reasoning as a sequence of temporal clue steps, each referencing a candidate time interval, instead of relying mainly on rewards for only the final prediction.",{"name":82,"@type":73,"acceptedAnswer":83},"What training signals does TimeThink use to improve learning?",{"text":84,"@type":76},"It introduces a step-wise temporal process reward for localized credit assignment and a joint process–outcome optimization objective to balance reasoning fidelity with task correctness.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]