[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83458-en":3,"doc-seo-83458-105":30,"detail-sidebar-cat-0-en-105":83},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83458,1099513958607,"Jiven","https://ap-avatar.wpscdn.com/avatar/100002390cf8733938c?x-image-process=image/resize,m_fixed,w_180,h_180&k=1778829742770036399",8,"Research & Report","Identifying and Resolving Pitfalls of Knowledge-Based VQA Benchmarks","Knowledge-Based Visual Question Answering (KB-VQA) evaluates whether visual language models can retrieve, ground, and reason over external structured knowledge beyond images, yet most benchmarks treat answer accuracy as a direct proxy for knowledge-grounded reasoning. This proxy depends on fragile assumptions: annotated answers must be derivable from the provided knowledge base, questions must be sufficiently constrained, and visuals must require grounded disambiguation. The paper shows systematic violations, including missing/contradicted answers, underspecified questions, and visually trivial single-entity scenes. It proposes an audit-and-repair protocol and a controlled multi-entity augmentation to restore derivability and challenge retrieval-based reasoning, leading to substantially revised performance trends and a call for verification-focused evaluation.","arXiv :2607 .00159v1 [ cs .CL] 30 Jun 2026  \nIdentifying and Resolving Pitfalls of Knowledge-Based VQA Benchmarks: Auditing, Repairing, and Augmenting  \nQian Ma 1, S M Rayeed 1, Charles V. Stewart 1, Qiong Wu2, and Yao  \nMa 1  \n1 Rensselaer Polytechnic Institute, Troy NY 12047, USA  \n2 AT&T Chief Data Office, Bedminster NJ 07921, USA  \n{maq5,rayees,stewart,[may13}@rpi.edu](may13}@rpi.edu), [qw6547@att.com](qw6547@att.com)  \nAbstract. Knowledge-Based Visual Question Answering (KB-VQA) aims to evaluate whether Visual Language Models (VLMs) can retrieve, ground, and reason over external structured knowledge beyond visual evidence.  \nIn practice, answer accuracy is widely adopted as the primary evaluation metric, implicitly treating correctness as a proxy for knowledgegrounded reasoning. However, for existing KB-VQA benchmarks, this proxy relies on critical assumptions that are often overlooked and rendered unreliable by benchmark issues: annotated answer must be derivable from the associated knowledge base, question must be well-posed with sufficient constraints, and visual setting must meaningfully require grounded disambiguation. In this work, we show that these assumptions are systematically violated in existing KB-VQA benchmarks. Our audit reveals substantial instances with missing or contradicted answersand underspecified questions that render accuracy a misleading metric.  \nFurthermore, we find that existing datasets rely on visually trivial, singleentity scenes that bypass the need for sophisticated visual-to-knowledge mapping. We demonstrate that even with controlled architectures, these flaws lead to distorted model rankings and overestimations of reasoning capabilities. To address this, we introduce (1) a principled audit-andrepair protocol that restores answer derivability and question clarity, and (2) a controlled multi-entity augmentation protocol that introduces visual ambiguity to challenge initial retrieval and grounded reasoning.  \nRe-evaluation under corrected and augmented settings yields markedly different performance trends. Our findings call for rethinking evaluation protocols and designing more interaction-aware KB-VQA benchmarks  \nthat prioritize verifiable reasoning over simple matching.  \nKeywords: KB-VQA · VLM · RAG  \n1  \n1 The datasets and code are available in [https://github.com/VAN-QIAN/ECCV26-ARA](https://github.com/VAN-QIAN/ECCV26-ARA). Work was initiated when Qian was an intern at AT&T CDO. Qiong Wu and Yao Ma are co-corresponding authors.  \n2 Q. Ma et al.  \n1 Introduction  \nRecent Vision-Language Models (VLMs) [1, 3, 29, 33] have demonstrated strong performance on a variety of visual reasoning tasks, especially Visual Question Answering (VQA) settings that require aligning image content with language understanding [2, 11, 15, 16] . Despite such effectiveness, they are still struggling to address tasks where answer is beyond the visual part like Knowledge-Based Visual Question Answering (KB-VQA) . KB-VQA [6, 22] is designed to evaluate whether a model can answer image-grounded questions correctly by retrieving and reasoning over an external, controlled knowledge base, rather than relying only on parametric memory [6, 9], where accuracy is used to assess such knowledge-grounded capability.  \nEmerging efforts are focusing on achieving better answer accuracy by developing re-ranker modules to select the most relevant and informative evidence segment for the answer generation [20,31], improving the model’s inherent capability to use all retrieved knowledge [7, 8, 13, 30] or empowering VLMs to invoke external tools to retrieve more relevant evidence [14] . Significantly increasing efforts and resources are being invested to achieve better accuracy score.  \nHowever, in representative KB-VQA benchmarks such as InfoSeek and EVQA [6,22], we observe recurring dataset issues (Section 3) that can make answer accuracy an unreliable indicator of the intended knowledge-grounded reasoning capability. 1. Answer-e","cbCaiokcEwbYtvjA","https://ap.wps.com/l/cbCaiokcEwbYtvjA","pdf",19877192,3,1,31,"English","en",105,"# Introduction\n## Answer–evidence misalignment\n## Underspecified questions and answer scope\n## Visually simplified scenes vs. multi-modal queries\n# Proposed audit-and-repair and augmentation approaches\n## Audit-and-repair protocol\n## Controlled multi-entity augmentation\n# Re-evaluation and implications","[{\"question\":\"How do the authors improve KB-VQA benchmarks?\",\"answer\":\"They introduce a principled audit-and-repair protocol to restore answer derivability and question clarity, and a controlled multi-entity augmentation protocol that adds visual ambiguity to better stress retrieval and grounded reasoning before evaluation.\"}]",1784188091,78,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":78,"head_meta":80,"extra_data":82,"updated_unix":28},"identifying-and-resolving-pitfalls-of-knowledge-based-vqa-benchmarks","",{"@graph":36,"@context":77},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/identifying-and-resolving-pitfalls-of-knowledge-based-vqa-benchmarks/83458/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71],{"name":72,"@type":73,"acceptedAnswer":74},"How do the authors improve KB-VQA benchmarks?","Question",{"text":75,"@type":76},"They introduce a principled audit-and-repair protocol to restore answer derivability and question clarity, and a controlled multi-entity augmentation protocol that adds visual ambiguity to better stress retrieval and grounded reasoning before evaluation.","Answer","https://schema.org",{"og:url":51,"og:type":79,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":81,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":84},[85,89,93,97,102,107,112,115,120,123,127],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":86,"show_sort_weight":87,"slug":88},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":90,"show_sort_weight":91,"slug":92},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Exam",70,"exam",{"id":98,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},5,"Comic",60,"comic",{"id":103,"doc_module":4,"doc_module_name":46,"category_name":104,"show_sort_weight":105,"slug":106},6,"Technology",50,"technology",{"id":108,"doc_module":4,"doc_module_name":46,"category_name":109,"show_sort_weight":110,"slug":111},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":113,"slug":114},30,"research-report",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},9,"Religion & Spirituality",20,"religion-spirituality",{"id":118,"doc_module":4,"doc_module_name":46,"category_name":121,"show_sort_weight":118,"slug":122},"World Cup","world-cup",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":124,"slug":126},10,"Lifestyle","lifestyle",{"id":128,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":98,"slug":130},19,"General","general"]