[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86255-en":3,"doc-seo-86255-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86255,687197207919,"Theodora","https://ap-avatar.wpscdn.com/avatar/a000253d6f5f7c60be?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779446848396160552",8,"Research & Report","See like a Robot: Robot-Centric Pointmaps for Vision-Language-Action Models","Vision-language-action (VLA) models translate visual observations and language instructions into robot actions, but actions are defined in the robot’s 3D coordinate frame while observations are typically captured in the camera frame. This frame mismatch is manageable with a fixed viewpoint yet becomes challenging when training aggregates demonstrations across diverse camera setups and requires cross-view generalization. The work introduces robot-centric pointmaps—images encoding per-pixel 3D robot-frame coordinates—preserving the dense H×W grid for pretrained 2D VLAs. Experiments on RoboCasa and real robots show improved performance, especially when camera placement differs from training.","arXiv :2607 . 11498v1 [ cs .RO] 13 Jul 2026  \nSee like a Robot: Robot-Centric Pointmaps for Vision-Language-Action Models  \nByungkun Lee 1∗ , Dongyoon Hwang 1∗ , Dongjin Kim 1 , Hojoon Lee2 , Minho Park 1 , Jaegul Choo 1  \n1 KAIST AI, 2Holiday Robotics  \nAbstract: Vision-language-action (VLA) models predict robot actions from visual observations and language instructions. These actions are defined in the robot’s own 3D coordinate frame, yet most VLAs observe the scene in the camera frame, creating a frame mismatch between where the scene is observed and where actions are defined. The mismatch is benign under a fixed viewpoint, where the policy can memorize a single observation-to-action mapping, but grows harder as large-scale datasets aggregate demonstrations across diverse camera setups and the policy must generalize this mapping across viewpoints. We address this mismatch with robot-centric pointmaps, images whose pixels store the 3D coordinates of scene points in the robot frame. Pointmaps provide robot-frame 3D geometry while preserving the dense H × W grid expected by pretrained 2D VLAs, so they integrate into existing VLAs with minimal architectural change. On RoboCasa, pointmaps improve both π0.5 and SmolVLA and outperform representative camera-viewpoint and 3D-aware baselines. In real-robot experiments, their advantage over an RGB-only policy widens when the camera is moved to a placement unseen during training. The project page is available here.  \nKeywords: VLA, manipulation, 3D geometry, pointmap  \n1 Introduction  \nVision-language-action (VLA) models learn robotic manipulation policies from large-scale datasets, predicting actions from visual observations and language instructions. These actions are commonly defined in the robot’s own 3D coordinate frame (e.g., task-space end-effector commands in the robot base frame), so reliable manipulation requires reasoning about where the target object lies in metric 3D coordinates relative to the robot. However, most VLAs receive observations in the camera frame, whether as RGB images or depth maps [1–5] . This creates a frame mismatch: the policy observes the scene in the camera frame but predicts actions in the robot frame.  \nThis frame mismatch becomes harder to handle when training data spans diverse camera viewpoints, as large-scale datasets increasingly aggregate demonstrations across institutions with different camera setups [6–8] . Under a fixed camera viewpoint, all observations share the same viewpoint relative to the robot, so a single observation-to-action mapping remains consistent across the dataset. Under a wide range of viewpoints, this no longer holds: the policy must additionally learn how observations from different viewpoints map to the robot-frame actions they should produce. The frame mismatch is present in both cases, but only under viewpoint variation must the policy generalize the observation-to-action mapping across viewpoints rather than memorize a single one.  \nThis motivates revising the policy’s input observation: rather than making the policy infer robotframe actions from camera-frame observations, we provide observations that already carry (1) 3D spatial information for manipulation and (2) that information in the robot’s own coordinate frame. As shown in Fig. 1, depth maps provide 3D cues, but remain tied to the camera frame rather than  \n*Equal [contribution. Correspondence to](contribution. Correspondence to byungkun.lee@kaist.ac.kr)[ byungkun.lee@kaist.ac.kr](contribution. Correspondence to byungkun.lee@kaist.ac.kr)  \nDiverse camera viewpoints in training data  \nDepth Pointcloud Pointmap  \n\n| Keeps 2D Grid? |  |  |  |\n| --- | --- | --- | --- |\n| Dense Observation? |  |  |  |\n| Robot-centric 3D Geometry? |  |  |  |\n\nFigure 1: Robot-centric pointmaps provide 3D geometry aligned with robot-frame actions. Large-scale training data is collected from diverse camera viewpoints around the robot. Robotcentric pointmaps preserve the dense H × W grid","cbCaijq8Rm2nhdsl","https://ap.wps.com/l/cbCaijq8Rm2nhdsl","pdf",17107143,5,1,20,"English","en",105,"# Introduction\n## Problem: frame mismatch in VLA training\n## Approach: robot-centric pointmaps\n## Integration with pretrained 2D VLAs","[{\"question\":\"What problem do robot-centric pointmaps address in vision-language-action (VLA) models?\",\"answer\":\"They address the frame mismatch between where observations are captured (camera frame) and where actions are defined (robot frame), which worsens under diverse camera viewpoints.\"},{\"question\":\"How does a pointmap represent robot-centric 3D information?\",\"answer\":\"A robot-centric pointmap stores, for each pixel, the corresponding scene point’s 3D coordinates expressed in the robot’s coordinate frame.\"},{\"question\":\"Why are pointmaps compatible with pretrained 2D VLA architectures?\",\"answer\":\"Pointmaps preserve the dense H×W image grid expected by pretrained 2D VLAs, enabling integration with minimal architectural changes rather than requiring point-cloud-specific processing.\"}]",1784209852,50,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"see-like-a-robot-robot-centric-pointmaps-for-vision-language-action-models","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/see-like-a-robot-robot-centric-pointmaps-for-vision-language-action-models/86255/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What problem do robot-centric pointmaps address in vision-language-action (VLA) models?","Question",{"text":76,"@type":77},"They address the frame mismatch between where observations are captured (camera frame) and where actions are defined (robot frame), which worsens under diverse camera viewpoints.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does a pointmap represent robot-centric 3D information?",{"text":81,"@type":77},"A robot-centric pointmap stores, for each pixel, the corresponding scene point’s 3D coordinates expressed in the robot’s coordinate frame.",{"name":83,"@type":74,"acceptedAnswer":84},"Why are pointmaps compatible with pretrained 2D VLA architectures?",{"text":85,"@type":77},"Pointmaps preserve the dense H×W image grid expected by pretrained 2D VLAs, enabling integration with minimal architectural changes rather than requiring point-cloud-specific processing.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,110,114,119,122,126,129,133],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":29,"slug":113},6,"Technology","technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":22,"slug":125},9,"Religion & Spirituality","religion-spirituality",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":127,"show_sort_weight":22,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":46,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":46,"category_name":135,"show_sort_weight":20,"slug":136},19,"General","general"]