[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85191-en":3,"doc-seo-85191-105":29,"detail-sidebar-cat-0-en-105":90},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},85191,962075114101,"Seraphina","https://ap-avatar.wpscdn.com/avatar/e000253a75eb197efd?x-image-process=image/resize,m_fixed,w_180,h_180&k=1780044092746381165",8,"Research & Report","What Does Your Short-Answer VQA Score Actually Measure Evaluator-Dependent Instability in Multimodal Short-Answer Benchmarks","Short-answer VQA benchmarks mix two different notions: semantic correctness of a model’s response and agreement with the surface form enforced by an automatic evaluator. The study examines this conflation across six vision–language models and six benchmarks using a human semantic judge (97.6% precision) to audit 37k+ official errors. A text-only judge reproduces the same benchmark-level false-negative pattern. On text-rich sets, up to half of flagged errors are semantically valid answers penalized for surface-form mismatch, with instability varying strongly by answer type. Prompt and context rewrites further flip item outcomes, and a deterministic CPU-only contract repair recovers part of the undercount, motivating semantic audits and answer-type diagnostics.","What Does Your Short-Answer VQA Score Actually Measure? Evaluator-Dependent Instability in Multimodal Short-Answer  \nBenchmarks  \n*  \nGuanhua Ye, Niu Jingbin, Yan Li, Meiyu Liang, Zhe Xue, Yingxia Shao, Yawen Li  \nBeijing University of Posts and Telecommunications  \nBeijing, China  \n{[g.ye](g.ye), kyungbin,liyanly, meiyu1210,xuezhe,shaoyx,[warmly0716}@bupt.edu.cn](warmly0716}@bupt.edu.cn)  \narXiv :2607 . 10240v1 [ cs .CV] 11 Jul 2026  \nAbstract  \nShort-answer VQA benchmarks conflate two distinct quantities: whether a model’s answer is semantically correct, and whether that answer matches the surface form expected by the automatic evaluator. We study this conflation across six vision–language models and six benchmarks, using a humanvalidated semantic judge (97.6% precision) to audit over 37k official errors. A second text-only judge reproduces the same benchmark-level false-negative pattern, showing that the effect is not an artifact of a single audit model. On text-rich benchmarks, up to half of these errors are semantically acceptable answers penalized purely for surface-form mismatch. This instability is structured by answer type: extractive and multi-span answers are far more evaluator-sensitive than scalar answers. Benign prompt and context rewrites further destabilize official outcomes, flipping item-level correctness at substantial rates without changing the underlying task. A deterministic CPU-only contract repair confirms that the undercount is partially recoverable. These findings imply that official short-answer VQA scores should be accompanied by semantic audits and answer-type diagnostics to remain interpretable.  \nCCS Concepts  \n• Computing methodologies → Natural language processing; Computer vision problems.  \nKeywords  \nvisual question answering, evaluation methodology, vision– language models, benchmark analysis, answer contracts  \n1 Introduction  \nA model looks at a photograph of a jet on a runway and is asked “What airline is this?” It outputs “Air France”—two words, correctly capitalized, unambiguously identifying the carrier. The gold answer is “Airfrance”—one word, no space. The benchmark marks it wrong. This is not a failure of visual understanding; the model read the livery correctly. It is a failure of the evaluation contract: the answer is semantically acceptable, but its surface form does not survive the stringmatching filter. Figure 1 shows this case alongside a second example where a highway-sign question is answered correctly in both form and content.  \n*Corresponding author.  \nFigure 1: Evaluator-sensitive undercount.  \nShort-answer VQA benchmarks have driven much of the progress in text-rich multimodal understanding over the past decade. ST-VQA [4] and TextVQA [32] stress scene text, DocVQA [27] and InfographicVQA [26] extend the challenge to documents and infographics, and ChartQA [25] adds visuallogical reasoning over charts. Despite these differences, they share one evaluation habit inherited from the original VQA setup [2]: compare the emitted string against one or more gold references under an automatic metric. Exact match, ANLS [29], and relaxed numeric accuracy differ in detail, but all make a string-level comparison decisive.  \nThat design fit an era of short, label-like outputs. It fits modern free-form generators less well. Once models answer in natural language rather than from a fixed vocabulary, answer wrapping (“The answer is 42”), morphology, and stylistic variation become part of the measured outcome. Ji et al. [14] argue that rigid exact-match patterns penalize correct but differently worded answers, and Ging et al. [10] likewise find substantial disagreement among exact-match, relaxed, and judge-based metrics on open-ended VQA.  \nThe consequence is that a leaderboard number—say, 84.3% on ST-VQA—is not pure question-answering success. It is the share of outputs that survive a benchmark-specific comparison rule applied to a benchmark-specific reference set. Some rejected outputs ","cbCaigXceS2HgzPg","https://ap.wps.com/l/cbCaigXceS2HgzPg","pdf",1019462,1,13,"English","en",105,"# Abstract\n# 1 Introduction\n## Evaluation contract and exact-match string comparison\n## Why benchmarks undercount semantic success\n## Answer-type dependence and prompt sensitivity\n# 2","[{\"question\":\"What two quantities do short-answer VQA benchmarks conflate?\",\"answer\":\"They conflate semantic correctness of the model’s answer with whether the answer’s surface form matches what the automatic evaluator expects under string-based matching rules.\"},{\"question\":\"How did the authors verify that false negatives are not caused by a single audit model?\",\"answer\":\"They used a human-validated semantic judge to audit over 37k official errors and then used a second, text-only judge that reproduced the same benchmark-level false-negative pattern.\"},{\"question\":\"Why do official short-answer VQA scores become unstable across benchmarks and answer types?\",\"answer\":\"The evaluator sensitivity depends on the implicit “answer contract” shaped by question and answer format (e.g., scalar vs extractive/multi-span vs OCR readouts), and benign prompt/context rewrites can flip question-level correctness without changing the underlying task.\"}]",1784201645,33,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":85,"head_meta":87,"extra_data":89,"updated_unix":27},"what-does-your-short-answer-vqa-score-actually-measure-evaluator-dependent-instability-in-multimodal-short-answer-benchmarks","",{"@graph":35,"@context":84},[36,53,67],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/what-does-your-short-answer-vqa-score-actually-measure-evaluator-dependent-instability-in-multimodal-short-answer-benchmarks/85191/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":61,"encodingFormat":60,"isAccessibleForFree":62,"interactionStatistic":63},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-16",true,{"@type":64,"interactionType":65,"userInteractionCount":4},"InteractionCounter",{"@type":66},"ViewAction",{"@type":68,"mainEntity":69},"FAQPage",[70,76,80],{"name":71,"@type":72,"acceptedAnswer":73},"What two quantities do short-answer VQA benchmarks conflate?","Question",{"text":74,"@type":75},"They conflate semantic correctness of the model’s answer with whether the answer’s surface form matches what the automatic evaluator expects under string-based matching rules.","Answer",{"name":77,"@type":72,"acceptedAnswer":78},"How did the authors verify that false negatives are not caused by a single audit model?",{"text":79,"@type":75},"They used a human-validated semantic judge to audit over 37k official errors and then used a second, text-only judge that reproduced the same benchmark-level false-negative pattern.",{"name":81,"@type":72,"acceptedAnswer":82},"Why do official short-answer VQA scores become unstable across benchmarks and answer types?",{"text":83,"@type":75},"The evaluator sensitivity depends on the implicit “answer contract” shaped by question and answer format (e.g., scalar vs extractive/multi-span vs OCR readouts), and benign prompt/context rewrites can flip question-level correctness without changing the underlying task.","https://schema.org",{"og:url":51,"og:type":86,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":88,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":91},[92,96,100,104,109,114,119,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":93,"show_sort_weight":94,"slug":95},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":97,"show_sort_weight":98,"slug":99},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":101,"show_sort_weight":102,"slug":103},"Exam",70,"exam",{"id":105,"doc_module":4,"doc_module_name":45,"category_name":106,"show_sort_weight":107,"slug":108},5,"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":45,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":45,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":45,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":45,"category_name":136,"show_sort_weight":105,"slug":137},19,"General","general"]