[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84575-en":3,"doc-seo-84575-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},84575,8796095360427,"Lucas Martin","https://ap-avatar.wpscdn.com/davatar_994ba38a5ba835b3df7d355c54d3ed8d",8,"Research & Report","MindEdit-Bench Object-Level Counterfactual Spatial Reasoning in VLMs","MindEdit-Bench benchmarks object-level counterfactual spatial reasoning for vision–language models (VLMs), moving beyond observational relation description. The benchmark builds six spatial tasks from three-photo smartphone triplets of newly captured indoor scenes using an automatic in-the-wild 3D scene-graph extraction pipeline. Four tasks evaluate perception and perspective transformation, while L4 (spatial editing) and L5 (cross-view visibility editing) require predicting outcomes after hypothetical object translation/rotation, with correct answers absent from all input images. Across 15 VLMs on 1,003 human-verified questions, mean accuracy is 8%–31% versus 81%–97% human majority vote, revealing non-uniform failure modes and large human–best-VLM gaps.","MindEdit-Bench: Benchmarking Object-Level Counterfactual Spatial Reasoning in VLMs from In-the-Wild Photos  \nLeyuan Yu 1 ,2 ,∗ , Xiao Tang 1 ,3 ,∗ , Minghao Liu 1 ,∗ , Xinyuan Li 1 ,2 Xiaokai Bai2 , Sheng Zhou2 , Qunshu Lin 1 ,†, Weihao Xuan4 , Naoto Yokoya4  \n1ZODA 2Zhejiang University 3Tongji University 4The University of Tokyo  \n∗Equal contribution. †Corresponding author.  \narXiv :2607 .0049 1v 1 [ cs .CV] 1 Jul 2026  \nAbstract  \nBenchmarks for vision–language models (VLMs) mostly test observational spatial reasoning: models describe relations already visible in the input. Existing what-if tasks typically vary the observer while keeping the scene fixed. Can VLMs instead predict the consequences of hypothetically moving or rotating an object? We introduce MindEdit-Bench1 , a benchmark of six spatial reasoning tasks built from three-photo smartphone triplets of newly captured indoor scenes via an automatic in-thewild 3D scene-graph extraction pipeline. Four tasks probe perception and perspective transformation over observed structure; two new tasks, L4 (spatial editing) and L5 (cross-view visibility editing), probe object-level counterfactual reasoning, where correct answers are absent from all input images. Each question provides 8–24 structured answer choices, enabling answer-letter-level diagnosis of spatial and fallback errors. The benchmark covers 120 private indoor scenes not drawn from public datasets, reducing public-data pretraining-overlap risk. Across 15 VLMs on 1,003 human-verified questions, task-wise mean VLM accuracy is only 8%–31%, versus 81%–97% human majorityvote accuracy. The pooled human–best-VLM gap is 53 pp, with at least 39 pp on every task. The structured answer space further reveals non-uniform failures, including weaker cameradepth-axis inference and fallback behavior on difficult visibility-editing cases.  \n1 Introduction  \nSpatial reasoning is central to cognitive science, computer vision, and embodied AI. In vision– language models (VLMs), spatial-reasoning evaluation has progressed from single-image QA to multiimage, video, and holistic benchmark suites (Chenet al., 2024 ; Cheng et al., 2024 ; Yang et al., 2025c ; Li et al., 2025 ; Yang et al., 2025a ; Wu et al., 2025 ;  \n1Dataset: [https://huggingface.co/datasets/](https://huggingface.co/datasets/)[ ](https://huggingface.co/datasets/)ZODAOfficial/MindEdit-Bench.  \nCai et al., 2025b ; Xiao et al., 2025) . Yet most existing benchmarks remain observational: they ask models to describe spatial relations already present in the input. This leaves a distinct question unanswered: can VLMs predict what follows when the spatial configuration itself is changed?  \nObject-level counterfactual spatial reasoning targets this missing case. It requires a model to represent a scene, apply a hypothetical intervention to an object, and infer the resulting spatial consequences. Existing VLM what-if benchmarks cover two narrower settings: observer-level tasks vary the viewer while keeping the scene fixed (Ma et al., 2023 ; Wang et al., 2025), and image-editing benchmarks modify pixels to evaluate generation fidelity rather than spatial reasoning (Xiao et al., 2026) . The object-level case, where an object is hypothetically translated or rotated while the observer remains fixed, has not been systematically evaluated.  \nWe introduce MindEdit-Bench, a six-task spatial reasoning benchmark built from three-photo smartphone triplets of privately captured indoor scenes. Four tasks test perception and perspective transformation over observed 3D structure (L1– L3b), while two test object-level counterfactual reasoning. L4 spatial editing asks how an object relation changes after virtual translation or rotation; L5 cross-view visibility editing asks which virtual edit would make an object appear or disappear in another view. Since L4/L5 answers appear in none of the input images, they remove direct 2D visual cues by construction and require reasoning over an internal 3D scene ","cbCaihu3ObWE3M0T","https://ap.wps.com/l/cbCaihu3ObWE3M0T","pdf",1409131,1,18,"English","en",105,"# Abstract\n# Introduction\n# Benchmark Overview\n## Spatial editing (L4)\n## Cross-view visibility editing (L5)\n## Evaluation setup","[{\"question\":\"MindEdit-Bench evaluates what capability in VLMs?\",\"answer\":\"It evaluates object-level counterfactual spatial reasoning—predicting spatial consequences when an object is hypothetically translated or rotated while the observer view stays fixed.\"},{\"question\":\"How are the tasks L4 and L5 different from the earlier perception/perspective tasks?\",\"answer\":\"L4 spatial editing and L5 cross-view visibility editing require object-level counterfactual answers that do not appear in any of the input photos, forcing reasoning over an internal 3D scene representation.\"},{\"question\":\"What is the dataset construction and evaluation scale reported?\",\"answer\":\"The benchmark uses 120 newly and privately captured indoor scenes, generating 1,003 human-verified questions. It evaluates 15 VLMs and uses structured multiple-choice options for answer-letter-level diagnostics.\"}]",1784196892,45,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"mindedit-bench-object-level-counterfactual-spatial-reasoning-in-vlms","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/mindedit-bench-object-level-counterfactual-spatial-reasoning-in-vlms/84575/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"MindEdit-Bench evaluates what capability in VLMs?","Question",{"text":75,"@type":76},"It evaluates object-level counterfactual spatial reasoning—predicting spatial consequences when an object is hypothetically translated or rotated while the observer view stays fixed.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How are the tasks L4 and L5 different from the earlier perception/perspective tasks?",{"text":80,"@type":76},"L4 spatial editing and L5 cross-view visibility editing require object-level counterfactual answers that do not appear in any of the input photos, forcing reasoning over an internal 3D scene representation.",{"name":82,"@type":73,"acceptedAnswer":83},"What is the dataset construction and evaluation scale reported?",{"text":84,"@type":76},"The benchmark uses 120 newly and privately captured indoor scenes, generating 1,003 human-verified questions. It evaluates 15 VLMs and uses structured multiple-choice options for answer-letter-level diagnostics.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]