[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85578-en":3,"doc-seo-85578-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85578,1649267921044,"Ava Thompson","https://us-avatar.wpscdn.com/avatar/1800007509477c92dfb?_k=1782875107921204101",8,"Research & Report","Towards Temporal Compositional Reasoning in Long-Form Sports Videos","Sports video analysis is a multimodal challenge because it requires understanding complex, dynamic human actions over long time spans. Even with rapid progress in multimodal large language models, long-horizon reasoning remains hard since answers demand locating and integrating temporally sparse evidence. The work attributes the gap to missing supervision for dispersed evidence and the absence of explicit temporal evidence localization and justification. It introduces SportsTime and proposes Chain-of-Time Reasoning (CoTR) for temporally grounded evidence composition.","arXiv :2604 .22226v2 [ cs .CV] 13 Jul 2026  \nTowards Temporal Compositional Reasoning in Long-Form Sports Videos  \nSiyu Cao 1 ,2, Lu ZhangB 1 ,2, Ruizhe Zeng 1 ,2, and Zhi-yong LiuB 1 ,2 ,3  \n1 MAIS, Institute of Automation, Chinese Academy of Sciences,  \nBeijing 100190, China  \n{caosiyu2024, lu.zhang, zengruizhe2022, [zhiyong.liu}@ia.ac.cn](zhiyong.liu}@ia.ac.cn)[ ](zhiyong.liu}@ia.ac.cn)2 School of Artificial Intelligence, University of Chinese Academy of Sciences,  \nBeijing 100049, China  \n3 Nanjing Artificial Intelligence Research of IA, Nanjing, 211100, China  \nAbstract. Sports video analysis is a challenging domain for multimodal understanding because it involves complex and dynamic human activities. Despite rapid progress in Multimodal Large Language Models (MLLMs), long-horizon reasoning in sports videos remains difficult, as answering questions requires both locating and integrating temporally sparse evidence into reasoning. We attribute this limitation to two closely related factors: insufficient supervision over temporally dispersed evidence and the lack of methods for explicit temporal evidence localization and justification. To address these gaps, we introduce SportsTime, a large-scale benchmark for long-form sports video understanding, comprising 14K+ open-ended QA pairs and 50K+ step-wise temporal evidence annotations. Building on SportsTime, we propose Chain-of-Time Reasoning (CoTR), which treats reasoning as a process of temporally grounded evidence composition. Specifically, during training, CoTR introduces a temporal-reward GRPO to encourage temporally grounded reasoning.  \nDuring inference, it employs an anchor-observe-infer evidence-seeking loop to iteratively localize, verify, and compose temporal evidence before producing the final answer. Experiments show that SportsTime exposes substantial gaps in current MLLMs, while CoTR yields consistent gains over strong baselines, improving both temporal compositional reasoning performance and step-wise grounding quality. The dataset and code are available at [https://github.com/ustiniansy/SportsTime](https://github.com/ustiniansy/SportsTime).  \nKeywords: Temporal Compositional Reasoning · Sports · Benchmark  \n1 Introduction  \nSports videos capture complex and dynamic human activities at scale, playing a central role in global cultural life and serving as a key data source for professional analytics. Over the past decade, artificial intelligence and computer vision techniques have fundamentally reshaped sports video analysis, enabling tasks such as  \n2 Cao et al.  \nFig. 1: Chain-of-Time reasoning enables more reliable and verifiable answers in long-form sports videos.  \nplayer detection and tracking, action spotting, tactical analysis, etc [7, 8, 12, 43] . More recently, the rapid development of Multimodal Large Language Models (MLLMs) [2, 6, 10] has opened new opportunities toward unified sports video understanding, particularly for more diverse and open-ended tasks such as commentary generation [27] and video question answering (VideoQA) [41], which require flexible and compositional inference over long-form multimodal content.  \nDespite significant advances, current MLLMs still struggle with long-horizon reasoning in videos [17, 18, 23, 28, 36, 54] . This weakness is particularly evident in sports scenarios, which are long-form and highly dynamic, where interpreting sparse but critical events demands a holistic and long-horizon understanding. For example, in a soccer match, understanding how a goal is scored may require tracing a sequence of events, such as passes and player movements that unfold long before the final shot. Such scenarios expose a key limitation of current MLLMs: they struggle to identify temporally dispersed evidence and compose it into reasoning, a capability we refer to as temporal compositional reasoning.  \nWe argue that this limitation mainly stems from two tightly coupled aspects:  \n(1) the scarcity of high-quality annotations that explic","cbCaimgNxXkH4ej9","https://ap.wps.com/l/cbCaimgNxXkH4ej9","pdf",2935430,3,1,18,"English","en",105,"# Introduction\n## Long-horizon reasoning challenges in sports videos\n## Temporal compositional reasoning and its causes\n## SportsTime benchmark and annotations\n## Chain-of-Time Reasoning (CoTR) approach","[{\"question\":\"Why is long-horizon reasoning in sports videos difficult for current multimodal large language models?\",\"answer\":\"Because answering requires locating and integrating temporally sparse evidence dispersed across many moments, and models lack explicit mechanisms to ground reasoning in that evidence.\"},{\"question\":\"What are the main factors the paper identifies for the limitation?\",\"answer\":\"Insufficient supervision over temporally dispersed evidence and the lack of methods for explicit temporal evidence localization and justification.\"},{\"question\":\"What does SportsTime provide, and how does CoTR use it?\",\"answer\":\"SportsTime offers 14K+ open-ended QA pairs and 50K+ step-wise temporal evidence annotations. CoTR treats reasoning as temporally grounded evidence composition, using a temporal-reward training strategy and an anchor-observe-infer evidence-seeking loop at inference.\"}]",1784204701,45,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"towards-temporal-compositional-reasoning-in-long-form-sports-videos","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/towards-temporal-compositional-reasoning-in-long-form-sports-videos/85578/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why is long-horizon reasoning in sports videos difficult for current multimodal large language models?","Question",{"text":75,"@type":76},"Because answering requires locating and integrating temporally sparse evidence dispersed across many moments, and models lack explicit mechanisms to ground reasoning in that evidence.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What are the main factors the paper identifies for the limitation?",{"text":80,"@type":76},"Insufficient supervision over temporally dispersed evidence and the lack of methods for explicit temporal evidence localization and justification.",{"name":82,"@type":73,"acceptedAnswer":83},"What does SportsTime provide, and how does CoTR use it?",{"text":84,"@type":76},"SportsTime offers 14K+ open-ended QA pairs and 50K+ step-wise temporal evidence annotations. CoTR treats reasoning as temporally grounded evidence composition, using a temporal-reward training strategy and an anchor-observe-infer evidence-seeking loop at inference.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]