[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82398-en":3,"doc-seo-82398-105":29,"detail-sidebar-cat-0-en-105":90},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},82398,1099514068365,"Aurelia","https://ap-avatar.wpscdn.com/avatar/10000253d8d9f28188e?_k=1776742907772140068",8,"Research & Report","The Count Is There, but Misaligned: Understanding and Correcting Counting Failures in VLMs","Despite strong performance on many multimodal tasks, vision-language models (VLMs) struggle with basic object counting. This work tests whether failures come from missing internal knowledge or from a mismatch between internal representations and verbalized outputs. Nonlinear probes trained on activations from four VLMs across five counting datasets reliably detect counting errors, showing that correct counts are often encoded even when answers are wrong. SVCCA reveals partially shared subspaces but misaligned readout directions, improved via causal steering.","The Count Is There, but Misaligned: Understanding and Correcting  \nCounting Failures in VLMs  \nAhmed Oumar El-Shangiti1,2 Abzal Nurgazy1 Hilal AlQuabeh1  \nNikolai Rozanov3 Kentaro Inui1  \n1MBZUAI 2DataBayt.AI Labs 3Imperial College London  \nCorrespondence: [ahmed.oumar@mbzuai.ac.ae](ahmed.oumar@mbzuai.ac.ae)  \narXiv :2607 .09544v1 [ cs .CV] 10 Jul 2026  \nAbstract  \nDespite strong performance on many multimodal tasks, vision-language models (VLMs) still struggle with basic object counting. We investigate whether this reflects missing internal knowledge or a gap between internal representations and verbalized outputs. Training simple probes on activations from four VLMs across five counting datasets reveals that nonlinear probes can reliably detect counting errors, suggesting that VLMs often encode the correct count even when they output the wrong answer. SVCCA analysis shows that probes trained on ground-truth counts and probes trained on model outputs occupy a partially shared activation subspace but read out along misaligned directions. We further validate our findings using a causal steering intervention, proving that strengthening the direction of count-identified probes does improve model counting performance. Motivated by this result, we propose a detector-guided self-correction method that selectively re-prompts the model only when an internal error detector predicts failure. This simple inference-time intervention improves counting accuracy by up to 15.6% absolute percentage points, without any parameter updates. Our results establish activation-based error probing as both a practical tool for improving VLM counting and a mechanistic lens on the gap between internal knowledge and model outputs.  \n1 Introduction  \nVision-language models (VLMs) have achieved strong performance across a wide range of tasks, including image captioning (Li et al., 2023a), visual question answering (Liu et al., 2023a), reasoning (Lu et al., 2024), and web navigation (Koh et al., 2024) . However, strong performance on these tasks does not imply reliable quantitative perception. Among such capabilities, counting is especially important because it requires a model to  \ndetect relevant objects, distinguish them from distractors, maintain consistent correspondences, and map visual evidence to an exact numerical answer. This makes counting a useful test of whether a VLM truly grounds its predictions in the image rather than relying on superficial correlations or language priors (Vo et al., 2025) .  \nAt the same time, counting exposes several core failure modes in VLMs, including hallucination, failures under clutter or occlusion, poor object individuation, and mistakes in translating perceptual representations into language. Recent work has shown that these weaknesses remain substantial even in strong contemporary models (Weng et al., 2025 ; Vo et al., 2025 ; Paiss et al., 2023a ; Fu et al., 2023) .  \nMost existing work documents that VLMs are weak at counting, but it does so primarily at the behavioral level, through benchmarks and aggregate performance comparisons (Weng et al., 2025 ; Vo et al., 2025 ; Paiss et al., 2023a ; Fu et al., 2023) . These studies are important because they establish counting as a persistent failure mode, yet they largely leave open a more fundamental question: why do VLMs fail at counting? In particular, behavioral evaluations can show when a model is wrong, but they do not reveal whether the correct count is absent from the model’s internal representations, whether it is present but poorly aligned with the model’s eventual answer, or whether the failure emerges later during decoding.  \nA smaller body of work begins to examine counting more mechanistically (Hasani et al., 2025 ; Alghisi et al., 2025), and some methods attempt to improve performance directly. For example, Alghisi et al. (2025) retrain parts of the model, while Sengupta et al. (2025) use attention-based interventions and obtain modest gains. However these","cbCaimb9cnWV8bcN","https://ap.wps.com/l/cbCaimb9cnWV8bcN","pdf",1819820,1,21,"English","en",105,"# Introduction\n## Counting as a Core Stress Test for VLMs\n## Behavioral Benchmarks vs Mechanistic Explanations\n## Multi-Probe Analysis of Counting Errors\n## SVCCA Alignment and Causal Steering\n## Detector-Guided Self-Correction Method","[{\"question\":\"What problem does the paper focus on regarding VLMs?\",\"answer\":\"The paper studies why vision-language models often fail at basic object counting, even when they perform well on many other multimodal tasks.\"},{\"question\":\"How do the authors detect whether counting is encoded internally?\",\"answer\":\"They train nonlinear probes on intermediate activations from four VLMs using supervision for ground-truth counts, the model’s output counts, and error likelihood, then test whether probe readouts reveal the correct count.\"},{\"question\":\"What does SVCCA show about counting representations?\",\"answer\":\"SVCCA indicates that probes trained with ground-truth versus output-related supervision occupy a partially shared activation subspace, but the signals are read out along misaligned directions.\"}]",1784180130,53,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":85,"head_meta":87,"extra_data":89,"updated_unix":27},"the-count-is-there-but-misaligned-understanding-and-correcting-counting-failures-in-vlms","",{"@graph":35,"@context":84},[36,53,67],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/the-count-is-there-but-misaligned-understanding-and-correcting-counting-failures-in-vlms/82398/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":61,"encodingFormat":60,"isAccessibleForFree":62,"interactionStatistic":63},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-16",true,{"@type":64,"interactionType":65,"userInteractionCount":4},"InteractionCounter",{"@type":66},"ViewAction",{"@type":68,"mainEntity":69},"FAQPage",[70,76,80],{"name":71,"@type":72,"acceptedAnswer":73},"What problem does the paper focus on regarding VLMs?","Question",{"text":74,"@type":75},"The paper studies why vision-language models often fail at basic object counting, even when they perform well on many other multimodal tasks.","Answer",{"name":77,"@type":72,"acceptedAnswer":78},"How do the authors detect whether counting is encoded internally?",{"text":79,"@type":75},"They train nonlinear probes on intermediate activations from four VLMs using supervision for ground-truth counts, the model’s output counts, and error likelihood, then test whether probe readouts reveal the correct count.",{"name":81,"@type":72,"acceptedAnswer":82},"What does SVCCA show about counting representations?",{"text":83,"@type":75},"SVCCA indicates that probes trained with ground-truth versus output-related supervision occupy a partially shared activation subspace, but the signals are read out along misaligned directions.","https://schema.org",{"og:url":51,"og:type":86,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":88,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":91},[92,96,100,104,109,114,119,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":93,"show_sort_weight":94,"slug":95},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":97,"show_sort_weight":98,"slug":99},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":101,"show_sort_weight":102,"slug":103},"Exam",70,"exam",{"id":105,"doc_module":4,"doc_module_name":45,"category_name":106,"show_sort_weight":107,"slug":108},5,"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":45,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":45,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":45,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":45,"category_name":136,"show_sort_weight":105,"slug":137},19,"General","general"]