[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-81626-en":3,"doc-seo-81626-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},81626,8796095360427,"Lucas Martin","https://ap-avatar.wpscdn.com/davatar_994ba38a5ba835b3df7d355c54d3ed8d",8,"Research & Report","Diagnosing Long Video Quantitative Reasoning in Multimodal LLMs via Enumeration and Counting","Final-answer video QA can judge whether a model outputs the correct count, but it often hides which instances were counted, when supporting evidence appeared, and why the model failed. This work diagnoses long-video quantitative reasoning in multimodal large language models through three coupled capabilities: enumerating query-relevant instances, temporally grounding supporting evidence spans, and aggregating evidence into final counts. EC-Bench provides 152 untrimmed 30+ minute videos with 1,699 open-ended queries and human-verified evidence spans, enabling fine-grained analysis.","Diagnosing Long-Video Quantitative Reasoning in Multimodal LLMs via  \nEnumeration and Counting  \nFumihiko Tsuchiya 1 Taiki Miyanishi 1 Shunsuke Yasuki 1 Mahiro Ukai2 Nakamasa Inoue2  \nShuhei Kurita3 Yusuke Iwasawa 1 Yutaka Matsuo 1  \n1The University of Tokyo, Japan  \n2Institute of Science Tokyo, Japan  \n3National Institute of Informatics, Japan  \narXiv :2603 .29943v2 [ cs .CV] 10 Jul 2026  \nAbstract  \nFinal-answer video QA can show whether a model predicts the right number, but not which instances it counted, when the supporting evidence occurs, or why it failed. We diagnose long-video quantitative reasoning in multimodal large language models (MLLMs) through three coupled abilities: enumerating query-relevant instances, temporally grounding supporting evidence, and aggregating the evidence into counts. To support this analysis, we build EC-Bench, an evidence-annotated evaluation suite with 152 untrimmed videos longer than 30 minutes, 1,699 open-ended queries across six reasoning categories, and human-verified evidence spans. We evaluate 22 open-source and proprietary MLLMs using timestamped visual frames and transcripts. The best average scores reach only 29.98% Enumeration F1 and 23. 74% Counting accuracy, compared with human performance of 78.57% and 82.97%, respectively. Our analyses show that counting errors are rarely isolated arithmetic mistakes: Enumeration F1 is strongly associated with Counting accuracy, temporal grounding quality is associated with lower counting error, and Counting accuracy drops as supporting evidence becomes more distributed. These findings recast long-video counting as evidence retrieval, temporal grounding, deduplication, and aggregation across the video, rather than simple numerical prediction.  \n1. Introduction  \nFinal-answer video question answering (QA) is limited asan evaluation of quantitative video understanding. A model may predict the correct number without revealing which instances it counted, when the supporting evidence occurred, or whether the answer was grounded in the video. This limitation becomes especially problematic in long-form videos, where relevant events can be sparse, visually diverse, and separated by tens of minutes. Recent long-video benchmarks  \nhave expanded evaluation to long-form and hour-scale settings [4, 44, 48, 54], but final-answer evaluation alone provides limited visibility into the evidence behind a count.  \nCounting offers a useful diagnostic setting because the answer is objective, yet obtaining it requires more than numerical prediction. A model must identify relevant instances, distinguish them from distractors, avoid duplicate counting across repeated appearances, and aggregate evidence under the constraints of the query. Broad video QA and MLLMoriented benchmarks include counting-related or temporal reasoning tasks [21, 23, 24, 28], while dedicated video counting datasets mainly study repetition, object/event counting, or audio-visual counting in short, trimmed, or temporally localized settings [9, 10, 30, 40, 62] . However, these evaluations typically do not diagnose which evidence supports the model’s count.  \nA central challenge in long-video counting is the gap between local evidence and global task structure. Individual evidence spans may be brief, but the model must search an untrimmed video, apply temporal or semantic constraints, deduplicate repeated or boundary-crossing instances, and produce a consistent count. Thus, long-video counting tests evidence retrieval, temporal grounding, deduplication, and aggregation over extended temporal contexts.  \nMotivated by this view, we formulate long-video quantitative reasoning through three coupled abilities. Enumeration requires listing query-relevant instances while avoiding irrelevant or duplicate items. Temporal grounding requires localizing the evidence spans that support the answer. Counting requires aggregating the identified evidence into a numerical answer. This formulation moves beyond fina","cbCaicxDeK8f1293","https://ap.wps.com/l/cbCaicxDeK8f1293","pdf",6918179,3,1,22,"English","en",105,"# Abstract\n# Introduction\n## Limitations of final-answer video QA\n## Counting as a diagnostic task\n## Challenge: local evidence vs global structure\n## Three coupled abilities\n# Contributions and evaluation overview","[{\"question\":\"What limitations does final-answer video QA have for quantitative counting?\",\"answer\":\"It can only indicate whether the predicted number is correct, without showing which instances were counted, when evidence occurred, or whether the answer was grounded in the video.\"},{\"question\":\"How does the paper diagnose long-video quantitative reasoning in multimodal LLMs?\",\"answer\":\"It evaluates three coupled abilities: enumerating query-relevant instances, temporally grounding supporting evidence spans, and aggregating the evidence into counts.\"},{\"question\":\"What is EC-Bench and what does it enable?\",\"answer\":\"EC-Bench is an evidence-annotated evaluation suite with 152 untrimmed 30+ minute videos and 1,699 open-ended queries, including human-verified evidence spans for detailed inspection of counting, grounding, and deduplication behavior.\"}]",1784174947,55,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"diagnosing-long-video-quantitative-reasoning-in-multimodal-llms-via-enumeration-and-counting","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/diagnosing-long-video-quantitative-reasoning-in-multimodal-llms-via-enumeration-and-counting/81626/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-26","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What limitations does final-answer video QA have for quantitative counting?","Question",{"text":75,"@type":76},"It can only indicate whether the predicted number is correct, without showing which instances were counted, when evidence occurred, or whether the answer was grounded in the video.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the paper diagnose long-video quantitative reasoning in multimodal LLMs?",{"text":80,"@type":76},"It evaluates three coupled abilities: enumerating query-relevant instances, temporally grounding supporting evidence spans, and aggregating the evidence into counts.",{"name":82,"@type":73,"acceptedAnswer":83},"What is EC-Bench and what does it enable?",{"text":84,"@type":76},"EC-Bench is an evidence-annotated evaluation suite with 152 untrimmed 30+ minute videos and 1,699 open-ended queries, including human-verified evidence spans for detailed inspection of counting, grounding, and deduplication behavior.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]