[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86541-en":3,"doc-seo-86541-105":30,"detail-sidebar-cat-0-en-105":83},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86541,962075114101,"Seraphina","https://ap-avatar.wpscdn.com/avatar/e000253a75eb197efd?x-image-process=image/resize,m_fixed,w_180,h_180&k=1780044092746381165",8,"Research & Report","GDP Benchmarking Grounded Multimodal Reasoning over Professional PDF Documents","Professional work in many domains relies on real PDF documents such as benefits packets, leases, datasheets, clinical guidelines, and construction plans. Existing document-AI benchmarks often evaluate capabilities separately in isolation, so high scores do not guarantee correct answers to realistic, field-specific questions grounded in the right evidence within a PDF. GDP.pdf introduces a 100-item benchmark with expert rubrics and capability tags, showing that top models achieve low pass rates and common failures include misaligned tables, chart reading errors, missed footnotes, scan noise, and amendment supersession.","GDP.pdf: Benchmarking Grounded Multimodal Reasoning over Professional PDF  \nDocuments  \nSuhaas Garre*, Emily Ritchie, Sushant Mehta, Edwin Chen  \nSurge AI  \narXiv :2607 . 11192v1 [ cs .CV] 13 Jul 2026  \nAbstract  \nA large share of day-to-day work in professional domains happens inside PDF files: benefits packets, leases, datasheets, clinical guidelines, construction plans. Benchmarks for document AI have generally measured the required capabilities in isolation: OCR, layout analysis, chart reasoning, table QA, document VQA. A high score on any oneof them does not necessarily reveal whether a model can answer a realistic question that someone in the field would actually ask about a specific PDF. GDP.pdf is a benchmark built to measure this directly. It consists of question–document pairs authored by working professionals in ten fields, and a candidate question was kept only when at least two frontier multimodal models failed it in a way that mattered: a wrong answer, missed decisive evidence, or a fabricated claim, rather than a superficial difference such as style. Each item comes with a rubric of atomic criteria, so we can report a graded rubric score as well as a strict task-level pass rate, and each item is tagged against a taxonomy of eleven capabilities in three tiers, spanning text extraction and grounding, table and chart comprehension, cross-referencing, spatial reasoning, and abstention on unsupported queries. We evaluated seven frontier models on the 100-item benchmark. The best model passed only 15% of the items and the worst passed 1%. Most errors trace back to a small set of recurring loss patterns: misaligned tables, misread charts, skipped footnotes and exclusions, miscounted floor-plan symbols, scan noise, and amendments that supersede earlier text. The full 100-item benchmark is publicly available at [https://](https://)[ ](https://)[huggingface.co/datasets/surgeai/GDP.pdf](huggingface.co/datasets/surgeai/GDP.pdf).  \n1. Introduction  \nMultimodal models are usually evaluated on visual QA of a fairly academic kind. The document tasks people would  \n* Correspondence to: [suhaas@surgehq.ai](suhaas@surgehq.ai)  \nAccepted at the 2nd Workshop on Knowledge-Intensive Multimodal Reasoning (KnowledgeMR) at CVPR 2026 . The workshop is non-archival.  \nactually like to hand to a model look different. Comparing plan tiers in a benefits packet, locating the indemnification clause in a lease, or counting fixtures on a floor plan all require working through a long and variably formatted file, and the fact that settles the answer may sit in a footnote three pages away from the flowchart it qualifies. Current models score well on the standard visual-reasoning suites, but we find those scores to be a poor guide to performance on the document workflows that power everyday economic activity.  \nSeveral separate problems are involved here, and prior benchmarks have tended to study each in isolation. The first is document structure (multi-page tables, sidebars, legends, footnotes, amendments appended at the end) . The second is background knowledge: a benefits table assumes, for example, that the reader knows what a “tier” is. The third problem, which was also a major motivation of this work, is that the failures are not visible as failures. The model cites a clause that exists and a number that is on the page; the clause is simply not the one that governs the user’s question.  \nGDP.pdf was constructed to evaluate all three problems jointly.1 The questions are phrased as real practitioners phrase them, the input is the original PDF rather than a cleaned-up extract, and the grading checks whether the response rests on the correct evidence. The current version contains 100 items covering ten domains (Finance, Healthcare, Legal, STEM/Research, Engineering, Construction, Manufacturing/Supply Chain, Insurance, Real Estate, and Human Resources) . Every item has an expert rubric and capability tags, and an item was admitted only after at leas","cbCaimyhRoJl07ij","https://ap.wps.com/l/cbCaimyhRoJl07ij","pdf",239503,3,1,9,"English","en",105,"# Introduction\n## Contributions\n# Related Work","[{\"question\":\"What are the common error patterns found in the pilot evaluation?\",\"answer\":\"Errors frequently stem from misaligned tables, misread charts, skipped footnotes/exclusions, miscounted floor-plan symbols, scan noise, and amendments that supersede earlier text.\"}]",1784212516,23,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":78,"head_meta":80,"extra_data":82,"updated_unix":28},"gdp-benchmarking-grounded-multimodal-reasoning-over-professional-pdf-documents","",{"@graph":36,"@context":77},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/gdp-benchmarking-grounded-multimodal-reasoning-over-professional-pdf-documents/86541/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71],{"name":72,"@type":73,"acceptedAnswer":74},"What are the common error patterns found in the pilot evaluation?","Question",{"text":75,"@type":76},"Errors frequently stem from misaligned tables, misread charts, skipped footnotes/exclusions, miscounted floor-plan symbols, scan noise, and amendments that supersede earlier text.","Answer","https://schema.org",{"og:url":51,"og:type":79,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":81,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":84},[85,89,93,97,102,107,112,115,119,122,126],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":86,"show_sort_weight":87,"slug":88},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":90,"show_sort_weight":91,"slug":92},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Exam",70,"exam",{"id":98,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},5,"Comic",60,"comic",{"id":103,"doc_module":4,"doc_module_name":46,"category_name":104,"show_sort_weight":105,"slug":106},6,"Technology",50,"technology",{"id":108,"doc_module":4,"doc_module_name":46,"category_name":109,"show_sort_weight":110,"slug":111},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":113,"slug":114},30,"research-report",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},"Religion & Spirituality",20,"religion-spirituality",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":120,"show_sort_weight":117,"slug":121},"World Cup","world-cup",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":123,"slug":125},10,"Lifestyle","lifestyle",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":98,"slug":129},19,"General","general"]