[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85905-en":3,"doc-seo-85905-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85905,7971461740886,"Theodore","https://ap-avatar.wpscdn.com/davatar_3d24733baf745e90a7e4bdd5f77d97b2",8,"Research & Report","SynthDocBench: Controlled Benchmark for Long-Context Visual Document Understanding","SynthDocBench introduces a fully synthetic, controlled benchmark for long-context visual document understanding, designed to attribute VLM failures to specific causes rather than confounded real-world factors. Using a combinatorial design, it independently varies document length, layout structure, modality composition, and question type, generating documents end to end via an LLM pipeline across multiple layout archetypes with random overrides to limit spurious correlations. Evaluations on seven frontier VLMs reveal three failure modes: degradation with length, positional sensitivity with a negative Early→Late trend, and breakdown of chart comprehension in long-document settings.","arXiv :2607 . 10400v1 [ cs .CV] 11 Jul 2026  \nSynthDocBench: Controlled Benchmark for LongContext Visual Document Understanding  \nAbhigya Verma1* , Khyati Mahajan1* , Amit Kumar Saha1* , Shruthan Radhakrishna1 , Sagar Davasam1 , Vikas Yadav1 , Sai Rajeswar1, 2, 3  \n1ServiceNow AI 2Mila 3Université de Montréal  \nVision language models (VLMs) have achieved strong performance on visual document understanding benchmarks such as DocVQA, ChartQA, and MMLongBenchDoc. However, real-world documents combine multiple factors such as length, layout complexity, modality, and question difficulty, which makes it difficult to attribute model failures to specific causes. We introduce SYNTHDOCBENCH, a fully synthetic benchmark for long-context visual document understanding that systematically controls factors including document length, layout structure, modality composition, and question type. The benchmark is constructed using a combinatorial design, each factor is varied independently across generated documents, enabling controlled analysis of model behavior. Documents are generated end to end using an LLM pipeline across six layout archetypes, with a 40 percent random override to prevent models from exploiting spurious correlations. Additionally, SynthDocBench spans long-context documents with substantially greater length and structural diversity than existing benchmarks. Evaluating seven frontier VLMs, we uncover three failure modes that existing benchmarks cannot surface: sharp degradation with document length, a systematic positional sensitivity in which the middle third of a document is hardest for five of six models and five of six models show a negative Early→Late trend (steepest decline: 8.3 percentage points), and breakdown of chart comprehension in long-document settings. These results suggest that current models may be overfitting to benchmark artifacts rather than achieving robust long-context visual document understanding.  \nCorrespondence: [abhigya.verma@servicenow.com](abhigya.verma@servicenow.com)  \nCode: [https://github.com/ServiceNow/SynthDocBench](https://github.com/ServiceNow/SynthDocBench)  \nDataset: [https://huggingface.co/datasets/ServiceNow-AI/](https://huggingface.co/datasets/ServiceNow-AI/)[ ](https://huggingface.co/datasets/ServiceNow-AI/)SynthDocBench  \n* Equal contribution.   \n1 Introduction  \nUnderstanding long, visually rich documents is a defining challenge for vision language models (VLMs) . Real-world documents interleave text, tables, charts, and complex layouts across dozens or hundreds of pages, demanding both long-range retrieval and cross-modal reasoning. Benchmarks such as DocVQA (Mathew et al., 2021), ChartQA (Masry et al., 2022), and MMLongBench-Doc (Ma et al., 2024) have driven substantial progress, yet this progress conceals a fundamental diagnostic blind spot: when a model fails on a real document, it is difficult to know why.  \nOn single-page tasks, evaluation is approaching saturation. Frontier models exceed 95% on DocVQA (Bai et al., 2025a; Wang et al., 2025a) and 89% on ChartQA (Anthropic, 2025; Bai et al., 2025a)—though harder benchmarks such as ChartQAPro (Masry et al., 2025), VisuLogic (Zhang et al., 2025), and ChartMuseum (Xu et al., 2025) show that chart understanding is far from solved, with degradations exceeding 30 percentage points. Yet all evaluate charts in isolation, abstracted from the multi-page contexts in which they naturally occur. Longcontext benchmarks like MMLongBench-Doc (Ma et al., 2024) address the document-length dimension, targeting long PDF comprehension with rich visual content, and the strongest model achieves only 57% . LongDocURL (Deng et al., 2025) and M-LongDoc (Chia et al.,  \n2025) extend to documents spanning 100s of pages, but prioritize breadth of coverage over controlled diagnosis: neither constructs questions requiring joint reasoning over charts and textually distributed evidence across distant pages. Furthermore, these benchmarks draw on real documents, c","cbCaihJlpu6AtbVP","https://ap.wps.com/l/cbCaihJlpu6AtbVP","pdf",2519749,3,1,35,"English","en",105,"# Introduction\n## Problem: diagnosing failures on real long documents\n## Related benchmarks and diagnostic blind spots\n## Synthetic benchmarks as a solution\n# SYNTHDOCBENCH: controlled benchmark design\n## Independent control axes and generation pipeline\n## Evaluation findings and failure modes","[{\"question\":\"What problem does SynthDocBench address in long-context visual document understanding?\",\"answer\":\"Real-world documents mix length, layout complexity, modality, and question difficulty, making it hard to determine why a model fails. SynthDocBench enables controlled diagnosis by varying factors independently.\"},{\"question\":\"How does SynthDocBench control document factors to make failures attributable?\",\"answer\":\"It uses a combinatorial design where document length, layout structure, modality composition, and question type are varied independently across generated documents, preventing reliance on spurious correlations.\"},{\"question\":\"What failure modes does the benchmark reveal when evaluating seven frontier VLMs?\",\"answer\":\"It finds three modes: sharp degradation as document length increases, systematic positional sensitivity with the middle third hardest and a negative Early→Late trend, and chart comprehension breakdown in long-document settings.\"}]",1784207070,88,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"synthdocbench-controlled-benchmark-for-long-context-visual-document-understanding","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/synthdocbench-controlled-benchmark-for-long-context-visual-document-understanding/85905/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-26","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does SynthDocBench address in long-context visual document understanding?","Question",{"text":75,"@type":76},"Real-world documents mix length, layout complexity, modality, and question difficulty, making it hard to determine why a model fails. SynthDocBench enables controlled diagnosis by varying factors independently.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does SynthDocBench control document factors to make failures attributable?",{"text":80,"@type":76},"It uses a combinatorial design where document length, layout structure, modality composition, and question type are varied independently across generated documents, preventing reliance on spurious correlations.",{"name":82,"@type":73,"acceptedAnswer":83},"What failure modes does the benchmark reveal when evaluating seven frontier VLMs?",{"text":84,"@type":76},"It finds three modes: sharp degradation as document length increases, systematic positional sensitivity with the middle third hardest and a negative Early→Late trend, and chart comprehension breakdown in long-document settings.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]