[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84930-en":3,"doc-seo-84930-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},84930,687197207639,"Asher","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","Lift3D VLA Lifting VLA Models to 3D Geometry and Dynamics-Aware Manipulation","Vision–Language–Action (VLA) models show broad task generalization, yet effective robotic manipulation requires accurate 3D geometry and spatial reasoning under changing dynamics. Existing 3D-aware VLA methods suffer from scarce training data and information loss from 3D encoding pipelines, preventing joint modeling of 3D structure and temporally structured actions. Lift3D-VLA introduces unified point-cloud reasoning, geometry-driven self-supervision, and layer-wise temporal action modeling to generate temporally coherent actions in dynamic environments.","Lift3D-VLA: Lifting VLA Models to 3D Geometry and Dynamics-Aware Manipulation  \nJiaming Liu†, Student Member, IEEE, Qingpo Wuwu†, Nuowei Han†, Hao Chen†, Zhuoyang Liu, Fan Fei, Yueru Jia, Chenyang Gu, Yandong Guo, Boxin Shi, Senior Member, IEEE, and Shanghang Zhang  \narXiv :2607 .06564v 1 [ cs .RO] 7 Jul 2026  \nAbstract—Recently, Vision–Language–Action (VLA) models have demonstrated strong generalization across diverse tasks. However, effective robotic manipulation in physical environments fundamentally requires geometric understanding and spatial reasoning. While some VLA approaches attempt to incorporate 3D information, they are constrained by limited data availability and geometric information loss in current 3D encoding pipelines, and fail to jointly capture 3D geometry and temporally structured actions in dynamic environments. To address these limitations, we introduce Lift3D-VLA, a unified VLA framework that equips models with explicit 3D point cloud reasoning and enables temporally coherent action generation. First, building upon our previous work Lift3D, an enhanced 2D model-lifting strategy is proposed to geometrically align 3D points with pretrained 2D positional embeddings. This design enables direct pointcloud encoding within the VLA vision encoder while minimizing spatial information loss. Based on explicit 3D inputs, we propose Geometry-Centric Masked Autoencoding (GC-MAE), a dualobjective self-supervised framework that reconstructs the current point cloud while predicting its future geometric evolution. This formulation allows the 2D vision encoder to internalize both 3D structure and physical dynamics. To fully exploit 3D representations, we further design layer-wise temporal action modeling, which leverages multiple layers of the LLM to collaboratively predict action chunks, enabling temporally consistent predictions. Across 22 simulated tasks and 8 real-world manipulation tasks, Lift3D-VLA achieves 10.8% and 11.1% higher mean success rates on MetaWorld and RLBench than the best-performing prior VLA methods, and outperforms the strongest real-world baseline by 4 percentage points, while exhibiting stronger generalization to out-of-distribution perturbations. Project website: [https://lift3dvla.github.io/](https://lift3dvla.github.io/).  \nIndex Terms—Robotic Manipulation, Vision–Language–Action Model, 3D Representation.  \nI. INTRODUCTION  \nVision–Language–Action (VLA) models have recently emerged as a promising paradigm for robotic manipulation [2]–[4], showing strong generalization across diverse tasks by combining visual perception, language conditioning, and action generation. Despite this progress, robotic manipulation fundamentally requires spatial reasoning in the physical world [5]–[9]: the robot must infer 3D structure, reason about geometric relationships (e.g., reachability, occlusion, and  \nJiaming Liu, Qingpo Wuwu, Nuowei Han, Zhuoyang Liu, Yueru Jia, Chenyang Gu, Fan Fei, Boxin Shi, and Shanghang Zhang are with the State Key Laboratory of Multimedia Information Processing and National Engineering Research Center of Visual Technology, School of Computer Science, Peking University, Beijing, China. Hao Chen is with the CUHK, Shatin, Hong Kong. Yandong Guo is with the AI2Robotics, Beijing, China.†Jiaming Liu, Qingpo Wuwu, Nuowei Han, and Hao Chen contributed equally as co-first authors. Corresponding author: Shanghang Zhang. E-mail:  \n{[jiamingliu@stu.pku.edu.cn](jiamingliu@stu.pku.edu.cn), [shanghang@pku.edu.cn](shanghang@pku.edu.cn})[}](shanghang@pku.edu.cn})  \ncontact), and plan actions that remain temporally consistent asthe geometry evolves. Purely 2D VLA pipelines often struggle to reliably capture these geometric constraints, particularly in cluttered or dynamic environments.  \nA natural direction is to explicitly inject 3D information into VLA models or manipulation policies, with existing approaches primarily falling into two paradigms, as shown in Figure 1 a) . First, some methods directly en","cbCaitVSSGTGzQkz","https://ap.wps.com/l/cbCaitVSSGTGzQkz","pdf",6549831,2,1,14,"English","en",105,"# Introduction\n## 3D limitations in existing VLA approaches\n## Lift3D prior work and its gaps\n## Lift3D-VLA contributions and framework overview","[{\"question\":\"Why do VLA models need explicit 3D geometry for robotic manipulation?\",\"answer\":\"Robotic manipulation in physical environments depends on spatial reasoning, including inferring 3D structure and geometric relations such as reachability and occlusion, while planning actions that remain consistent as geometry evolves.\"},{\"question\":\"What are the main limitations of current 3D-aware VLA approaches?\",\"answer\":\"They are constrained by limited 3D data and often lose geometric fidelity through lossy 3D encoding pipelines or cross-modal transformations, which weakens structural correspondence and limits scalability.\"},{\"question\":\"How does Lift3D-VLA address geometry and temporal action challenges?\",\"answer\":\"It uses an enhanced 2D model-lifting strategy to align 3D points with pretrained 2D positional embeddings, applies Geometry-Centric Masked Autoencoding to reconstruct current point clouds and predict future geometric evolution, and adds layer-wise temporal action modeling to produce temporally consistent action chunks.\"}]",1784199418,35,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"lift3d-vla-lifting-vla-models-to-3d-geometry-and-dynamics-aware-manipulation","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/lift3d-vla-lifting-vla-models-to-3d-geometry-and-dynamics-aware-manipulation/84930/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-23","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why do VLA models need explicit 3D geometry for robotic manipulation?","Question",{"text":75,"@type":76},"Robotic manipulation in physical environments depends on spatial reasoning, including inferring 3D structure and geometric relations such as reachability and occlusion, while planning actions that remain consistent as geometry evolves.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What are the main limitations of current 3D-aware VLA approaches?",{"text":80,"@type":76},"They are constrained by limited 3D data and often lose geometric fidelity through lossy 3D encoding pipelines or cross-modal transformations, which weakens structural correspondence and limits scalability.",{"name":82,"@type":73,"acceptedAnswer":83},"How does Lift3D-VLA address geometry and temporal action challenges?",{"text":84,"@type":76},"It uses an enhanced 2D model-lifting strategy to align 3D points with pretrained 2D positional embeddings, applies Geometry-Centric Masked Autoencoding to reconstruct current point clouds and predict future geometric evolution, and adds layer-wise temporal action modeling to produce temporally consistent action chunks.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]