[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83097-en":3,"doc-seo-83097-105":29,"detail-sidebar-cat-0-en-105":83},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},83097,1099514067415,"Rowan","https://ap-avatar.wpscdn.com/avatar/100002539d78ffe74a7?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779092875211072502",8,"Research & Report","HoloCount: A Holistic Visual Counting Benchmark for MLLMs","Visual counting is a core pillar of multimodal intelligence, demanding precise fine-grained grounding and reliable spatial reasoning. While Multimodal Large Language Models (MLLMs) perform strongly on qualitative scene understanding, they often show critical quantitative weaknesses, including persistent numerical hallucinations. Existing counting benchmarks emphasize simplified perception and miss failure patterns under logical constraints and adversarial conditions. HoloCount introduces a diagnostic, three-level taxonomy covering Semantic, Analytical, and Robustness evaluations, and reports a major performance gap across 20+ state-of-the-art MLLMs as tasks progress to deeper reasoning.","HoloCount: A Holistic Visual Counting Benchmark  \nfor MLLMs  \nJinhong Deng  \nMeituan  \n[dengjinhong@meituan.com](dengjinhong@meituan.com)  \nLimeng Qiao  \nMeituan  \n[qiaolm@pku.edu.cn](qiaolm@pku.edu.cn)  \nGuanglu Wan  \nMeituan  \n[wanguanglu@meituan.com](wanguanglu@meituan.com)  \narXiv :2607 .06420v 1 [ cs .CV] 7 Jul 2026  \nAbstract  \nVisual counting is a fundamental pillar of multimodal intelligence, requiring a seamless integration of fine-grained grounding and spatial reasoning. While Multimodal Large Language Models (MLLMs) have achieved remarkable success in qualitative scene understanding, their quantitative precision remains a significant bottleneck, often characterized by persistent numerical hallucinations. Existing counting benchmarks primarily focus on basic perception in simplified contexts, failing to capture the complex failure modes that emerge under logical constraints or adversarial conditions. To address these limitations, we introduce HoloCount, a holistic and diagnostically rich benchmark structured around a three-level hierarchical taxonomy. HoloCount evaluates MLLMs across: (1) Semantic Counting, focusing on atomic and property-based enumeration; (2) Analytical Counting, assessing logical composition through spatial and set-based reasoning; and (3) Robustness Testing, probing model integrity against adverse scenarios and grounded counter-priors, such as high-density scenes and linguistic biases. Through an exhaustive evaluation of over 20 state-of-the-art MLLMs, we reveal a critical performance gap: even top-tier models degrade significantly as tasks transition from perception to complex analytical reasoning and adverse scenarios. Our findings provide a systematic landscape of current MLLM counting capabilities and offera roadmap for developing more grounded and reliable multimodal systems. The dataset is available at [https://mm-mvr.github.io/HoloCount/](https://mm-mvr.github.io/HoloCount/) .  \n1 Introduction  \nVisual counting [35, 29, 8], the ability to quantify instances of specific categories in an image, is an important skill that reflects a key aspect of visual intelligence. For modern Multimodal Large Language Models (MLLMs) [19, 4, 43, 39, 40, 12], counting is not merely a numerical task but a rigorous test of fine-grained grounding [20, 47] and spatial reasoning [7] . While MLLMs excel at qualitative scene understanding such as Visual Question Answering (VQA) [51], their limited quantitative precision remains a critical barrier for applications such as logistics [45], safety monitoring [24], and retail analytics [1], where numerical errors carry high costs.  \nDespite the rapid evolution of MLLMs [19, 4, 43, 39] such as the GPT series [3] and the QwenVL series [4, 43], their ability to perform reliable counting remains surprisingly fragile [29, 35] . Although these models exhibit strong zero-shot performance in general scene description, they often struggle with precision when tasked with enumerating objects. This counting limitation frequently manifests as numerical hallucinations, where the model confidently reports a count that bears little resemblance to the true quantity, undermining reliability in count-sensitive environments. Recent benchmarking efforts [29, 8, 35], including CountBench [29] and CountQA [35], have begun to address this challenge but remain constrained by limited object diversity and a focus on simple perceptual tasks. They often fail to reveal complex failure modes that arise when models face property constraints, logical constraints, or adversarial visual conditions.  \n| Model Acc |  |  |\n| --- | --- | --- |\n|  | Qwen3.5-397B-A17B | 76.9 |\n|  | Qwen3.5-27B | 75.9 |\n|  | Qwen3.5-122B-A10B | 75.4 |\n| Qwen3.5-35B-A3B |  | 75.0 |\n| Gemini-3-Flash-Preview |  | 74.8 |\n| Gemini-3.1-Pro-Preview |  | 74.7 |\n| Kimi-K2.6 |  | 73.0 |\n| Qwen3.5-9B |  | 71.4 |\n| Kimi-K2.5 |  | 68.8 |\n| Qwen3.5-4B |  | 68.6 |\n| Gemini-2.5-Pro |  | 67.9 |\n| GPT-5.5 |  | 67.6 |\n| Qwen3.5-27B-Instruct |  | 6","cbCaitLiTJTQJlh0","https://ap.wps.com/l/cbCaitLiTJTQJlh0","pdf",14428194,1,21,"English","en",105,"# Introduction\n## Visual counting and its importance for MLLMs\n## Limitations of existing counting benchmarks\n## HoloCount overview and taxonomy","[{\"question\":\"How does HoloCount evaluate MLLMs differently?\",\"answer\":\"HoloCount uses a three-level hierarchical taxonomy and evaluates across Semantic Counting, Analytical Counting, and Robustness Testing. This spans atomic and property-based enumeration, logical composition via spatial and set reasoning, and robustness against adverse scenarios and counter-priors.\"}]",1784185211,53,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":78,"head_meta":80,"extra_data":82,"updated_unix":27},"holocount-a-holistic-visual-counting-benchmark-for-mllms","",{"@graph":35,"@context":77},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/holocount-a-holistic-visual-counting-benchmark-for-mllms/83097/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71],{"name":72,"@type":73,"acceptedAnswer":74},"How does HoloCount evaluate MLLMs differently?","Question",{"text":75,"@type":76},"HoloCount uses a three-level hierarchical taxonomy and evaluates across Semantic Counting, Analytical Counting, and Robustness Testing. This spans atomic and property-based enumeration, logical composition via spatial and set reasoning, and robustness against adverse scenarios and counter-priors.","Answer","https://schema.org",{"og:url":51,"og:type":79,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":81,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":84},[85,89,93,97,102,107,112,115,120,123,127],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":86,"show_sort_weight":87,"slug":88},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":90,"show_sort_weight":91,"slug":92},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Exam",70,"exam",{"id":98,"doc_module":4,"doc_module_name":45,"category_name":99,"show_sort_weight":100,"slug":101},5,"Comic",60,"comic",{"id":103,"doc_module":4,"doc_module_name":45,"category_name":104,"show_sort_weight":105,"slug":106},6,"Technology",50,"technology",{"id":108,"doc_module":4,"doc_module_name":45,"category_name":109,"show_sort_weight":110,"slug":111},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":113,"slug":114},30,"research-report",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},9,"Religion & Spirituality",20,"religion-spirituality",{"id":118,"doc_module":4,"doc_module_name":45,"category_name":121,"show_sort_weight":118,"slug":122},"World Cup","world-cup",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":124,"slug":126},10,"Lifestyle","lifestyle",{"id":128,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":98,"slug":130},19,"General","general"]