[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82175-en":3,"doc-seo-82175-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},82175,687197207057,"Sage","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","OmniMapBench: 基于多样地图文档的视觉中心推理基准评测","OmniMapBench addresses a key weakness in many document-understanding benchmarks: visual content can be converted into text, letting models succeed without true visual grounding. The benchmark provides 2,096 manually annotated question–answer pairs over 1,603 map documents from nine categories, targeting skills from perception to multi-step visual reasoning. A benchmark-level metric, the Visual Dependency Index (VDI), measures accuracy drop when images are replaced by question-agnostic descriptions. Evaluations of 25 leading LVLMs show a large performance gap, highlighting the difficulty and validating OmniMapBench’s visual-centric focus.","arXiv :2607 .09068v1 [ cs .CV] 10 Jul 2026  \nOmniMapBench: Benchmarking Visual-Centric Reasoning on Diverse Map Documents  \nYang Chen 1 ,2∗, Yunwen Li3 ∗ , Yufan Shen 1 ∗ , Minghao Liu3 ,4 ∗ , Tianyu Zheng3 , Bin Fu 1 , Qunshu Lin4 , Zhi Yu2 , and Botian Shi 1 ,5 B  \n1 Shanghai Artificial Intelligence Laboratory 2 Zhejiang University  \n3 M-A-P 4 2077AI 5 Shanghai Innovation Institute [zjucheny@gmai.com](zjucheny@gmai.com)  \nAbstract. Recent advancements in LVLMs necessitate robust benchmarks for complex, visually grounded reasoning. A critical limitation is identified in many document understanding benchmarks: visual content is often reducible to text, enabling high performance without genuine visual grounding. To address this limitation, OmniMapBench is introduced to foster visual-centric reasoning for map documents. The benchmark comprises 2,096 manually annotated question-answer pairs across 1,603 map documents from nine categories. It is designed to probe a hierarchy of skills, ranging from perception to multi-step visual reasoning. To quantify benchmark properties, a simple yet effective benchmark-level metric is proposed: the Visual Dependency Index (VDI), defined as the accuracy drop when images are replaced with question-agnostic descriptions. OmniMapBench exhibits higher VDI than established benchmarks, which quantitatively validates its focus on irreducible visual reasoning.  \nComprehensive evaluations of 25 leading LVLMs are conducted on OmniMapBench. A significant performance gap is observed, with the topperforming model achieving only 75 .03% accuracy. This result underscores the challenges posed by OmniMapBench to current LVLMs. This work aims to catalyze progress in visual-centric reasoning for document understanding of LVLMs. The dataset and code are publicly available at [https://github.com/SIGMME/OmniMapBench](https://github.com/SIGMME/OmniMapBench).  \nKeywords: Visual reasoning · Map understanding · MLLM  \n1 Introduction  \nAdvances in the reasoning capabilities of Large Language Models (LLMs) [17,48, 69] have catalyzed the development of Large Vision-Language Models (LVLMs)[4,8,20,51,56] . Consequently, visual reasoning has emerged as a central research area [27, 66, 72, 88, 89] . The objective is shifting from perception tasks, such as object recognition and captioning, towards complex inference grounded in visual content. To this end, various “think with images” [52,64,65,90] methods are  \nexplored, wherein models are trained to decompose intricate visual queries and ∗ Equal contribution. B Corresponding author.  \n2 Y. Chen et al.  \n(a) Case from the DocVQA  \nQ: What percentage of smokers feel the need to find more excitement and sensation in life? A: 70  \n(b) Case from the OmniMapBench  \nFig. 1: Comparison of textualizability across benchmarks. (a) A DocVQA case with structured tabular content is easily described in text, and the question is solvable from the description. (b) In contrast, the simple map image from OmniMapBench cannot be sufficiently captured by a question-agnostic textual description to solve the question. The task requires visual-centric reasoning that is irreducible to a linear text format.  \nexecute programmatic actions via external tools [13, 77] . This paradigm facilitates a more adaptive and compositional approach to problem-solving, enabling models to address challenges that transcend simple visual perception.  \nThe rapid evolution of model capabilities, however, places new demands on evaluation benchmarks. As shown in Figure 1, a significant limitation is identified in many existing document Visual Question Answering (VQA) benchmarks [46]: the visual content in these benchmarks is often reducible to structured text. For instance, document images can be converted into markdown representations through OCR and layout analysis [7, 10, 19, 32, 39, 63], which is a common practice in retrieval-augmented generation (RAG) pipelines. Charts and tables are frequently expressible as code or d","cbCaioTRxxaKmZbE","https://ap.wps.com/l/cbCaioTRxxaKmZbE","pdf",2939840,2,1,20,"English","en",105,"# Introduction\n## Benchmark motivation and limitation of existing document VQA\n## Maps as an irreducible visual document category\n## OmniMapBench overview and benchmark construction\n## Evaluation setup and results","[{\"question\":\"What problem does OmniMapBench target in existing document understanding benchmarks?\",\"answer\":\"It targets the limitation that many benchmarks allow visual content to be reduced to text, so models can achieve high scores without genuine visual grounding.\"},{\"question\":\"How is OmniMapBench constructed and what does it cover?\",\"answer\":\"OmniMapBench contains 2,096 manually annotated question–answer pairs across 1,603 map documents spanning nine categories, designed to test perception to multi-step visual reasoning.\"},{\"question\":\"What is the Visual Dependency Index (VDI) and what does it measure?\",\"answer\":\"VDI measures accuracy drop when images are replaced with question-agnostic textual descriptions, quantifying how much the benchmark depends on irreducible visual information.\"}]",1784178595,50,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"omnimapbench-benchmarking-visual-centric-reasoning-on-diverse-map-documents","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/omnimapbench-benchmarking-visual-centric-reasoning-on-diverse-map-documents/82175/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-21","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does OmniMapBench target in existing document understanding benchmarks?","Question",{"text":75,"@type":76},"It targets the limitation that many benchmarks allow visual content to be reduced to text, so models can achieve high scores without genuine visual grounding.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How is OmniMapBench constructed and what does it cover?",{"text":80,"@type":76},"OmniMapBench contains 2,096 manually annotated question–answer pairs across 1,603 map documents spanning nine categories, designed to test perception to multi-step visual reasoning.",{"name":82,"@type":73,"acceptedAnswer":83},"What is the Visual Dependency Index (VDI) and what does it measure?",{"text":84,"@type":76},"VDI measures accuracy drop when images are replaced with question-agnostic textual descriptions, quantifying how much the benchmark depends on irreducible visual information.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,114,119,122,126,129,133],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":29,"slug":113},6,"Technology","technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":22,"slug":125},9,"Religion & Spirituality","religion-spirituality",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":127,"show_sort_weight":22,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":46,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":46,"category_name":135,"show_sort_weight":106,"slug":136},19,"General","general"]