[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-81905-en":3,"doc-seo-81905-105":31,"detail-sidebar-cat-0-en-105":93},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":30},81905,8796095462418,"Noah","https://ap-avatar.wpscdn.com/avatar/80000253c1241d02b47?x-image-process=image/resize,m_fixed,w_180,h_180&k=1778826106357471780",8,"Research & Report","Does It Fail to See or Fail to Know? Attributing Errors in Vision-Language Models","Vision-language models (VLMs) excel in visual question answering when required evidence is clearly visible, yet they struggle when questions demand knowledge not directly observable. Effective uncertainty quantification must both estimate failure likelihood and explain the underlying cause, including perception, entity recognition, and knowledge retrieval. This work proposes a unified framework to disentangle these failure modes and tests whether pre-generation signals predict them. Experiments across datasets and model families show two major error sources—visual/recognition bottlenecks and post-recognition factual uncertainty—enabling targeted interventions such as image repair, recognition support, or external retrieval.","Does It Fail to See or Fail to Know? Attributing Errors in Vision-Language  \nModels  \nKhang Nhat Hoang Vo1 Artem Vazhentsev1 Artem Shelmanov1 Timothy Baldwin1,2 Yova Kementchedjhieva1  \n1MBZUAI 2The University of Melbourne  \n[Correspondence:](Correspondence: Khang.Vo@mbzuai.ac.ae)[ Khang.Vo@mbzuai.ac.ae](Correspondence: Khang.Vo@mbzuai.ac.ae) [Yova.Kementchedjhieva@mbzuai.ac.ae](Yova.Kementchedjhieva@mbzuai.ac.ae)  \narXiv :2607 .04683v2 [ cs .CV] 10 Jul 2026  \nAbstract  \nVision-language models (VLMs) perform well on visual question answering with high-quality images but struggle when questions require knowledge beyond what is clearly and directly visible. In such settings, uncertainty quantification should not only indicate whether the model is likely to fail but also diagnose why it is uncertain, across dimensions such as perception, entity recognition, and knowledge retrieval. While prior work has focused on individual failure modes in isolation or treated incorrect answers as monolithic failures, we propose a unified framework for disentangling these failure modes and investigate whether pre-generation signals can predict these failure sources. Across a range of datasets and model families, we find a consistent pattern in VLM errors: some failures arise from visual or recognition bottlenecks, while others persist after the relevant entity is identified. Our main finding is that these failure sources can be predicted before decoding: recognition-related failures are best captured by visual-token representations, while failures that remain after recognition are better captured by prompt-conditioned hidden states. This pre-generation signal enables efficient failure-source prediction before the model produces an answer, allowing uncertain cases to be routed to targeted interventions such as image repair, entity recognition support, or external retrieval.  \n1 Introduction  \nModern vision-language models (VLMs) excel in visual perception over high-quality natural images (Li et al., 2023a ; Liu et al., 2023) . However, users of assistive technologies, educational tools, search systems, and digital assistants often ask questions that go beyond what is directly visible, such as Where is this actor from? or What is the natural habitat of this plant? This form of knowledge-intensive visual-question answering  \nPale Tussock Moth, native to Asia and Europe  \nMukden Palace, Shenyang, China  \nKnown entity, unknown fact  \n\n|  |  | Factual check |  |  |\n| --- | --- | --- | --- | --- |\n|  | What area is this insect native to?\u003Cbr>The East Coast of North America |  |  |  |\n| \u003Cbr>Visual recognition check |  |  |  |  |\n|  | \u003Cbr>Is this a Pale Tussock Moth?\u003Cbr>\u003Cbr>Yes |  |  |  |\n\nUnknown Entity  \n\n|  |  | Factual check |  |\n| --- | --- | --- | --- |\n|  | In what country is this complex located?\u003Cbr>In Beijing \u003Cbr> |  |  |\n| \u003Cbr>Visual recognition check\u003Cbr> |  |  |  |\n\nFigure 1: Examples of two attribution outcomes in knowledge-intensive VQA. For each image, we evaluate the target VLM with two independent checks: afactual VQA question and an entity-recognition probe. Top: the model answers the factual question incorrectly but recognizes the entity, so the error is attributed to UNKNOWN FACT. Bottom: the model answers the factual question incorrectly and also fails the recognition probe, so the error is attributed to UNKNOWN ENTITY.  \n(Schwenk et al., 2022 ; Marino et al., 2019 ; Chenet al., 2023) involves multi-step reasoning: from visual recognition, through entity linking, fact retrieval, and finally, answer generation, with any of these steps being susceptible to failure. Fig. 1 illustrates two distinct paths to an incorrect answer. Prior work has focused on individual failure modes in isolation or treated incorrect answers as monolithic failures. Some studies have proposed targeted probes and tailored mitigation strategies, including image-quality assessment and repair for  \ndegraded inputs (Cai et al., 2025a), evaluations of distributional ga","cbCainuwRB2F8Nx9","https://ap.wps.com/l/cbCainuwRB2F8Nx9","pdf",558556,6,1,22,"English","en",105,"# Abstract\n# Introduction\n## Knowledge-intensive visual question answering\n## Prior failure-mode work and limitations\n## Proposed goal and pre-generation decision-wise view\n## Paper contributions","[{\"question\":\"Why do vision-language models fail on knowledge-intensive visual question answering?\",\"answer\":\"They perform well when answers rely on visible evidence, but fail when questions require knowledge beyond what is directly observable. The failure can originate from perception, entity recognition, or knowledge retrieval steps.\"},{\"question\":\"How does the paper recommend diagnosing model uncertainty?\",\"answer\":\"It argues that uncertainty quantification should not only predict whether the model will fail, but also identify why it is uncertain by attributing the failure source across different dimensions such as recognition and factual retrieval.\"},{\"question\":\"What pre-generation signals can predict different failure sources?\",\"answer\":\"Recognition-related failures are best captured by visual-token representations, while failures that remain after entity recognition are better captured by prompt-conditioned hidden states. These signals allow predicting error sources before decoding the answer.\"}]","Does It Fail to See or Fail to Know? Attributing Errors in Vision-Language Models | PDF",1784176980,55,{"code":4,"msg":32,"data":33},"ok",{"site_id":25,"language":24,"slug":34,"title":13,"keywords":35,"description":14,"schema_data":36,"social_meta":88,"head_meta":90,"extra_data":92,"updated_unix":29},"does-it-fail-to-see-or-fail-to-know-attributing-errors-in-vision-language-models","",{"@graph":37,"@context":87},[38,55,70],{"@type":39,"itemListElement":40},"BreadcrumbList",[41,45,49,52],{"item":42,"name":43,"@type":44,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":46,"name":47,"@type":44,"position":48},"https://docshare.wps.com/document/","Document",2,{"item":50,"name":12,"@type":44,"position":51},"https://docshare.wps.com/document/research-report/",3,{"item":53,"name":13,"@type":44,"position":54},"https://docshare.wps.com/document/does-it-fail-to-see-or-fail-to-know-attributing-errors-in-vision-language-models/81905/",4,{"url":53,"name":13,"@type":56,"author":57,"headline":13,"publisher":59,"fileFormat":62,"inLanguage":24,"description":14,"dateModified":63,"datePublished":64,"encodingFormat":62,"isAccessibleForFree":65,"interactionStatistic":66},"DigitalDocument",{"name":9,"@type":58},"Person",{"url":42,"name":60,"@type":61},"DocShare","Organization","application/pdf","2026-08-03","2026-07-16",true,{"@type":67,"interactionType":68,"userInteractionCount":20},"InteractionCounter",{"@type":69},"ViewAction",{"@type":71,"mainEntity":72},"FAQPage",[73,79,83],{"name":74,"@type":75,"acceptedAnswer":76},"Why do vision-language models fail on knowledge-intensive visual question answering?","Question",{"text":77,"@type":78},"They perform well when answers rely on visible evidence, but fail when questions require knowledge beyond what is directly observable. The failure can originate from perception, entity recognition, or knowledge retrieval steps.","Answer",{"name":80,"@type":75,"acceptedAnswer":81},"How does the paper recommend diagnosing model uncertainty?",{"text":82,"@type":78},"It argues that uncertainty quantification should not only predict whether the model will fail, but also identify why it is uncertain by attributing the failure source across different dimensions such as recognition and factual retrieval.",{"name":84,"@type":75,"acceptedAnswer":85},"What pre-generation signals can predict different failure sources?",{"text":86,"@type":78},"Recognition-related failures are best captured by visual-token representations, while failures that remain after entity recognition are better captured by prompt-conditioned hidden states. These signals allow predicting error sources before decoding the answer.","https://schema.org",{"og:url":53,"og:type":89,"og:title":13,"og:site_name":60,"og:description":14},"article",{"robots":91,"canonical":53},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":94},[95,99,103,107,112,116,121,124,129,132,136],{"id":21,"doc_module":4,"doc_module_name":47,"category_name":96,"show_sort_weight":97,"slug":98},"Story & Novel",90,"story-novel",{"id":48,"doc_module":4,"doc_module_name":47,"category_name":100,"show_sort_weight":101,"slug":102},"Literature",80,"literature",{"id":54,"doc_module":4,"doc_module_name":47,"category_name":104,"show_sort_weight":105,"slug":106},"Exam",70,"exam",{"id":108,"doc_module":4,"doc_module_name":47,"category_name":109,"show_sort_weight":110,"slug":111},5,"Comic",60,"comic",{"id":20,"doc_module":4,"doc_module_name":47,"category_name":113,"show_sort_weight":114,"slug":115},"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":47,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":47,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":47,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":47,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":47,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":47,"category_name":138,"show_sort_weight":108,"slug":139},19,"General","general"]