[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86338-en":3,"doc-seo-86338-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86338,7971461741311,"Ophelia","https://ap-avatar.wpscdn.com/avatar/74000253aff267980c6?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779345379180704826",8,"Research & Report","Beyond the Single Camera: Agentic Multi-View Reasoning in Sports Video Understanding","Recent Multimodal Large Language Models (MLLMs) excel on single-view sports video understanding, but sports events often feature dense occlusion, rapid motion, and complex interactions that cannot be resolved from one camera. Existing benchmarks rarely test multi-view reasoning, even though real matches use multiple angles to provide complementary evidence for officiating. This work introduces SportMV-Bench with 787 multi-view video bundles and 2592 QA pairs across PAR, REI, and ADR, plus SportMV-Agent for iterative view selection, perception execution, and evidence-grounded reasoning, yielding 14.46% relative improvement.","Beyond the Single Camera: Agentic Multi-View Reasoning in Sports Video Understanding  \nKerui Chen 1 ,∗ , Jinglu Wang2 , Xiaoyi Zhang2 , Yan Lu2 ,†  \n1Zhejiang University 2Microsoft Research Asia  \narXiv :2607 . 11844v1 [ cs .CV] 13 Jul 2026  \nAbstract  \nRecent Multimodal Large Language Models (MLLMs) achieve strong performance on single-view video understanding benchmarks. However, sports videos involve dense occlusion, rapid motion, and complex interactions that are difficult to resolve from a single viewpoint. In practice, sports events are recorded from multiple camera angles, providing complementary evidence used by referees. Yet, no existing benchmark evaluates MLLMs on multi-view sports video understanding. To address this gap, we introduce SportMV-Bench, a comprehensive benchmark built from official match recordings, through a dedicated pipeline combining LLM-based generation, MLLM-based verification, and human filtering to ensure quality and consistency. SportMV-Bench containing 787 multi-view video bundlesand 2592 question-answer pairs across three categories: Perception-Aware Recognition (PAR), Rule-aware Event Interpretation (REI), and Adjudicative Decision Reasoning (ADR). Our analysis shows that current MLLMs fail to effectively exploit multi-view information, with the bottlenecks lying in fine-grained visual perception and view selection rather than logical reasoning or domain knowledge. We propose SportMV-Agent, an agentic framework that orchestrates an iterative loop of active view selection, perception tool execution, and evidence-grounded reasoning, achieving a significant 14.46% relative improvement over the strongest MLLM baseline.  \n1. Introduction  \nSports video understanding has received significant attention due to its wide applications in analytics [28], coaching [12], officiating support [15, 16], content retrieval [44], and audience engagement [29] . Recent Multimodal Large Language Models (MLLMs) [1, 2, 5, 21, 36] demonstrate strong capabilities in video understanding and reasoning,  \n†Corresponding author.  \n∗ This work was done when K. Chen was an intern at Microsoft Research Asia.  \nmaking them promising for sports video analysis. However, sports videos pose unique challenges for MLLMs, such as rapid movements, dense player interactions, domainspecific rules, and the need for fine-grained subtle visual cues. Several benchmarks have been developed to study sports video understanding, such as SoccerNet [11], SportsQA [22], and multi-sport datasets [41, 42] . While these benchmarks advance holistic understanding of broadcast videos, they typically assume that a single camera view provides sufficient evidence for question answering (QA) .  \nIn real-world, sports events are captured from multiple viewpoints to provide a complete perspective for refereesand audiences. Occlusion is common in sports: players often cluster around decisive moments, making critical actions hard to observe from a single view. Moreover, athletes may intentionally obscure events through exaggerated contact or simulated falls, making single-view evidence not only incomplete but sometimes misleading. As illustrated in Fig. 1(a), using only View 1 or View 2 leads to an incorrect out-of-bounds decision because the decisive touch is occluded, while incorporating View 3 reveals the correct outcome. Beyond occlusion, some queries require combining complementary evidence across views that may even appear conflicting. For example, in Fig. 1(b), View 1 localizes contact inside the penalty area, while View 2 shows that the defender did not touch the ball. Only by combining both views can the correct decision be made: a yellow card to attacker, and award opponents a defensive free kick rather than a penalty. These examples highlight the importance of multi-view reasoning for sports video understanding. Professional officiating systems such as VAR, Instant Replay, and Hawk-Eye already rely on multi-camera setups for this reason. However, e","cbCaieKQxERmRjZm","https://ap.wps.com/l/cbCaieKQxERmRjZm","pdf",4484910,6,1,14,"English","en",105,"# Introduction\n## SportMV-Bench benchmark design\n## Multi-view reasoning challenges in sports\n## Experiments and key findings\n## SportMV-Agent method overview","[{\"question\":\"Why single-view sports video understanding is insufficient for MLLMs?\",\"answer\":\"Sports videos involve dense occlusion, rapid motion, and subtle visual cues. Single-camera evidence can be incomplete or even misleading, requiring evidence aggregation across views.\"},{\"question\":\"What is SportMV-Bench and what does it contain?\",\"answer\":\"SportMV-Bench is a multi-view sports video benchmark built from official match recordings. It includes 787 multi-view video bundles and 2592 QA pairs across perception-aware recognition, rule-aware event interpretation, and adjudicative decision reasoning.\"},{\"question\":\"How does SportMV-Agent improve performance over strong MLLM baselines?\",\"answer\":\"SportMV-Agent orchestrates an iterative loop of active view selection, perception tool execution, and evidence-grounded reasoning. It achieves a 14.46% relative improvement over the strongest MLLM baseline by better exploiting decisive multi-view evidence.\"}]",1784210553,35,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"beyond-the-single-camera-agentic-multi-view-reasoning-in-sports-video-understanding","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/beyond-the-single-camera-agentic-multi-view-reasoning-in-sports-video-understanding/86338/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"Why single-view sports video understanding is insufficient for MLLMs?","Question",{"text":76,"@type":77},"Sports videos involve dense occlusion, rapid motion, and subtle visual cues. Single-camera evidence can be incomplete or even misleading, requiring evidence aggregation across views.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"What is SportMV-Bench and what does it contain?",{"text":81,"@type":77},"SportMV-Bench is a multi-view sports video benchmark built from official match recordings. It includes 787 multi-view video bundles and 2592 QA pairs across perception-aware recognition, rule-aware event interpretation, and adjudicative decision reasoning.",{"name":83,"@type":74,"acceptedAnswer":84},"How does SportMV-Agent improve performance over strong MLLM baselines?",{"text":85,"@type":77},"SportMV-Agent orchestrates an iterative loop of active view selection, perception tool execution, and evidence-grounded reasoning. It achieves a 14.46% relative improvement over the strongest MLLM baseline by better exploiting decisive multi-view evidence.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,111,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":107,"slug":138},19,"General","general"]