[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-81717-en":3,"doc-seo-81717-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},81717,3848291630094,"Emma Wilson","https://eur-avatar.wpscdn.com/davatar_085a072bc5b1113ac321206ff7593b45",8,"Research & Report","EmbodimentSemantic Spatial Scene-Graph Dataset and Benchmark for Vision-Language Models","Spatial grounding remains a key limitation of vision–language–action (VLA) systems for robotic manipulation, where models can follow language yet lack explicit representations of spatial arrangement, including support, containment, ordering, occlusion, and depth-sensitive relations. EMBODIMENTSEMANTIC introduces a dataset and benchmark that encodes scenes as directed object–relation–object triplets over a fixed relation set. It provides real-world observations with a SO101 arm plus simulator-grounded LIBERO benchmarks. The work evaluates relational grounding and tests whether scene graphs improve VLA control via structured prompt injection, revealing failures in depth-aware, viewpoint-dependent structure.","arXiv :2607 .00020v 1 [ cs .RO] 6 Jun 2026  \nEmbodimentSemantic: A Spatial Scene-Graph Dataset and Benchmark for Vision-Language Models on Embodied Manipulation Trajectories  \nHassan Jaber 1 , Refinath S N2 , Luca Cagliero 1 ,  \nChristopher E. Mower2 , Haitham Bou-Ammar2 ,3  \n1Politecnico di Torino, Italy,  \n2Huawei, Noah’s Ark Lab, United Kingdom,  \n3University College London, United Kingdom  \nAbstract: Spatial grounding remains a key limitation of vision–language–action (VLA) systems for robotic manipulation. While current models can recognize objects and follow language instructions, they often lack an explicit representation of how objects are arranged in space, including support, containment, ordering, occlusion, and depth-sensitive relations. We introduce EMBODIMENTSEMANTIC, a spatial scene-graph dataset and benchmark for evaluating relational grounding in embodied manipulation. EMBODIMENTSEMANTIC represents scenes as directed object–relation–object triplets, where each triplet specifies a spatial relation between an ordered pair of objects using a fixed set of relations. This representation enables direct evaluation of object binding, relation prediction, and spatial consistency. The dataset includes real-world manipulation observations collected with the low-cost SO101 robot arm, together with generated scene graphs for studying spatial grounding in practical robotic settings. To provide controlled validation, we also introduce a simulator-grounded LIBERO benchmark with over 60K manipulation frames and more than 120K camera-specific scene graphs across paired third-person and wrist views, where ground-truth relations are derived automatically from MuJoCo geometry, world coordinates, camera projections, and visibility constraints. We further test whether scene graphs improve downstream control by injecting them into existing VLA policy prompts. Experiments across open-source and commercial VLMs show that current models often predict plausible relations but struggle with exact depth-aware and viewpoint-dependent spatial structure.  \nEMBODIMENTSEMANTIC provides a unified framework for diagnosing spatial grounding in VLM perception and testing its utility for VLA manipulation.  \n1 Introduction  \nSince the development of the transformer architecture [1], robot learning has seen several major advances [2, 3] . Recent vision–language–action (VLA) models extend this trend by adapting internetscale vision–language models (VLMs) to robot control, typically through fine-tuning on robot demonstrations [4, 5] . Despite their scale, however, these models remain weakly grounded in scene geometry. Several works show that VLA policies can learn shortcut associations between actions and task-irrelevant visual or semantic context, such as background, texture, viewpoint, familiar objects, or linguistic priors, rather than reliably representing spatial relations such as distance, relative pose, object size, and viewpoint [6, 7, 8] .  \nRecent work has begun to address this gap by injecting spatial supervision and structured spatial representations into VLMs and VLAs, including large-scale metric spatial VQA data, robot-relevant annotations from 3D scans, and explicit egocentric spatial encodings or action grids [9, 10, 11] . However, these methods still largely rely on learned spatial priors and do not remove the fundamental  \nFigure 1: LIBERO scene-graph comparison. The top row shows simulator-grounded ground-truth relations, and the bottom row shows gemini-3.1-pro predictions for the same frames. Green edges indicate correct triplets and red edges indicate errors, highlighting both successful relation recovery and remaining depth-sensitive failures.  \nambiguity of inferring metric geometry from RGB appearance alone. For example, if one object has a larger image footprint than another, it may be physically larger, closer to the camera, or both. This size–distance ambiguity means that appearance alone does not uniquely determine metric scal","cbCaihKmkDdAqAZa","https://ap.wps.com/l/cbCaihKmkDdAqAZa","pdf",5052285,4,1,22,"English","en",105,"# Introduction\n## Spatial grounding limitations in VLA systems\n## Prior approaches and remaining ambiguity\n## Evaluation challenges for structured scene recovery\n## Dataset and benchmark contributions\n## Method and experimental evaluation","[{\"question\":\"What problem does EMBODIMENTSEMANTIC address in vision–language–action systems?\",\"answer\":\"It targets the weakness in spatial grounding for robotic manipulation, where models often lack an explicit representation of how objects are arranged in space, including depth-sensitive spatial relations.\"},{\"question\":\"How does EMBODIMENTSEMANTIC represent scenes?\",\"answer\":\"It represents each scene as directed object–relation–object triplets, where each triplet specifies a spatial relation between an ordered pair of objects using a fixed set of relations.\"},{\"question\":\"How are the benchmarks validated and what data sources are used?\",\"answer\":\"EMBODIMENTSEMANTIC includes real-world manipulation observations collected with a low-cost SO101 robot arm, and it also introduces a simulator-grounded LIBERO benchmark with relations derived from MuJoCo geometry, camera projections, coordinates, and visibility constraints.\"}]",1784175600,55,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"embodimentsemantic-spatial-scene-graph-dataset-and-benchmark-for-vision-language-models","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/embodimentsemantic-spatial-scene-graph-dataset-and-benchmark-for-vision-language-models/81717/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does EMBODIMENTSEMANTIC address in vision–language–action systems?","Question",{"text":75,"@type":76},"It targets the weakness in spatial grounding for robotic manipulation, where models often lack an explicit representation of how objects are arranged in space, including depth-sensitive spatial relations.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does EMBODIMENTSEMANTIC represent scenes?",{"text":80,"@type":76},"It represents each scene as directed object–relation–object triplets, where each triplet specifies a spatial relation between an ordered pair of objects using a fixed set of relations.",{"name":82,"@type":73,"acceptedAnswer":83},"How are the benchmarks validated and what data sources are used?",{"text":84,"@type":76},"EMBODIMENTSEMANTIC includes real-world manipulation observations collected with a low-cost SO101 robot arm, and it also introduces a simulator-grounded LIBERO benchmark with relations derived from MuJoCo geometry, camera projections, coordinates, and visibility constraints.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]