[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82121-en":3,"doc-seo-82121-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},82121,1099514067415,"Rowan","https://ap-avatar.wpscdn.com/avatar/100002539d78ffe74a7?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779092875211072502",8,"Research & Report","MultiView-Bench: A Diagnostic Benchmark for World-Centric Multi-View Integration in VLMs","Recent benchmarks for vision-language models (VLMs) evaluate single- or limited-view perception, leaving the key cognitive ability of world-centric multi-view integration insufficiently tested. MultiView-Bench is introduced as a diagnostic benchmark for holistic 3D scene comprehension, requiring models to separate object placement from transient perspectives and ground it in a fixed global coordinate system. Evaluations of frontier VLMs reveal strong performance on 2D planar relations, but consistent difficulty with 3D spatial relations and cross-view aggregation. View biases also emerge, including sensitivity to axis direction and texture variations, and a ViewNavigator multi-agent framework is proposed to actively select and fuse informative views, improving performance under strict budgets.","arXiv :2607 .08970v 1 [ cs .CV] 9 Jul 2026  \n.  \nMultiView-Bench: A Diagnostic Benchmark for World-Centric Multi-View Integration in VLMs  \nHantao Zhang 1 ,† Jinru Sui2 Ed Li 1  \nDirk Bergemann 1 Zhuoran Yang 1  \n1Yale University 2University of Edinburgh  \nAbstract. Recent benchmarks for VLMs largely assess single-or limited-view perception, leaving untested the core cognitive ability to integrate observations across viewpoints into a coherent, worldcentric (allocentric) 3D mental model. We introduce MultiView-Bench, a diagnostic benchmark expressly designed to evaluate multi-view integration for holistic 3D scene comprehension. Unlike existing datasets that focus on pixel-level mapping or camera-relative navigation, MultiView-Bench requires models to decouple object positioning from transient perspectives and ground them in a fixed global coordinate system. This capability serves as a prerequisite for VLMs before being deployed for downstream tasks such as mechanical part assembly. Our systematic evaluation of frontier VLMs reveals consistent failure modes: strong performance on 2D planar relations from a single image, but marked difficulty with 3D spatial relations and with aggregating information across views. We further identify biases in VLMs, such as struggles with unconventional axis directions and sensitivity to object colorways and texture variations. Acknowledging these limitations, we propose ViewNavigator, a multi-agent framework that actively selects informative viewpoints, perceives, and fuses multi-view evidence, improving diverse base models on MultiView-Bench even under a strict budget-matched comparison (and by 3-5× for the full agent) .  \n[Project Webpage:](Project Webpage: hantaozhangrichard.github.io/MultiView-Bench)[ hantaozhangrichard.github.io/MultiView-Bench](Project Webpage: hantaozhangrichard.github.io/MultiView-Bench)  \nIntroduction  \nRecent advances in Large Language Models (LLMs) (Brown et al., 2020; Achiam et al., 2023; Hurst et al., 2024; Touvron et al., 2023) and Vision-Language Models (VLMs) (Radford et al., 2021; Li et al., 2022; Google, 2023; Dai et al., 2023) have demonstrated remarkable progress in complex perceptual and reasoning tasks, including spatial navigation (Zhou et al., 2024; Yamada et al., 2023) and image understanding (Dosovitskiy, 2020) . Their strong generalization capabilities, coupled with emergent reasoning skills, make them compelling candidates for cognitive systems that integrate perception and strategic planning (Bubeck et al., 2023) . When equipped with appropriate tools and scaffolding, such systems have shown promise in robotics control (Zitkovich et al., 2023), 3D modeling (Hu et al., 2024; Gu et al., 2025), and image editing (Huang et al., 2024) .  \nHowever, effectively solving many of these tasks fundamentally depends on the ability to perceive and reason about scenes from multiple viewpoints (Edelman, 1998; Bülthoff and Edelman, 1992) . Humans naturally perform multi-angle observations to construct coherent mental models of objects, resolving perceptual ambiguities  \n† [Correspondence:](Correspondence: {hantao.zhang}@yale.edu)[ {hantao.zhang}@yale.edu](Correspondence: {hantao.zhang}@yale.edu)  \nthat arise from single viewpoints (Shepard and Metzler, 1971) . This ability is crucial when assembling complex objects, where each component must be rotated and inspected from multiple viewpoints to determine how it connects with others. In contrast, a single static image often fails to convey critical structural or relational details necessary for accurate reasoning and manipulation, underscoring the importance of multi-view perception in spatial cognition (Marr, 2010) .  \nA possible workaround involves geometric representations such as point clouds, meshes, or voxels (Qi et al., 2017a; Wu et al., 2015; Mescheder et al., 2019), which encode precise 3D coordinates and shapes. However, processing such low-level geometric data typically requires specialized encoders (Qi et","cbCaiaIz3uekpr95","https://ap.wps.com/l/cbCaiaIz3uekpr95","pdf",10956045,2,1,33,"English","en",105,"# Introduction\n## Problem motivation: multi-view integration vs. existing benchmarks\n## Proposed solution: MultiView-Bench\n## Contributions and evaluation setup","[{\"question\":\"What core ability does MultiView-Bench diagnose in VLMs?\",\"answer\":\"It evaluates a VLM’s ability to integrate observations across multiple viewpoints into a coherent, world-centric 3D mental model grounded in a fixed global coordinate system.\"},{\"question\":\"How does MultiView-Bench differ from existing multi-view benchmarks?\",\"answer\":\"Existing benchmarks mainly test egocentric reasoning such as perspective-taking, camera-motion effects, or view-dependent navigation, while MultiView-Bench targets view-invariant, holistic 3D grounding and cross-view aggregation.\"},{\"question\":\"What failures and biases are reported for frontier VLMs on MultiView-Bench?\",\"answer\":\"Models often perform well on 2D planar relations from a single image but struggle with 3D spatial relations and combining information across views, showing biases such as difficulty with unconventional axis directions and sensitivity to object color and texture variations.\"}]",1784178324,83,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"multiview-bench-a-diagnostic-benchmark-for-world-centric-multi-view-integration-in-vlms","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/multiview-bench-a-diagnostic-benchmark-for-world-centric-multi-view-integration-in-vlms/82121/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-19","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What core ability does MultiView-Bench diagnose in VLMs?","Question",{"text":75,"@type":76},"It evaluates a VLM’s ability to integrate observations across multiple viewpoints into a coherent, world-centric 3D mental model grounded in a fixed global coordinate system.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does MultiView-Bench differ from existing multi-view benchmarks?",{"text":80,"@type":76},"Existing benchmarks mainly test egocentric reasoning such as perspective-taking, camera-motion effects, or view-dependent navigation, while MultiView-Bench targets view-invariant, holistic 3D grounding and cross-view aggregation.",{"name":82,"@type":73,"acceptedAnswer":83},"What failures and biases are reported for frontier VLMs on MultiView-Bench?",{"text":84,"@type":76},"Models often perform well on 2D planar relations from a single image but struggle with 3D spatial relations and combining information across views, showing biases such as difficulty with unconventional axis directions and sensitivity to object color and texture variations.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]