[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82848-en":3,"doc-seo-82848-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},82848,2336464648322,"Aria","https://ap-avatar.wpscdn.com/avatar/2200025388227c56fec?_k=1778556882303663488",8,"Research & Report","Geometry-Aware Motion Latents for Learning Robust Manipulation Policies","Geometry-Aware Motion Latents (GeoMoLa) learns discrete motion latent codes for robotic manipulation by predicting how 3D point clouds evolve during actions. Instead of reconstructing visual observations, GeoMoLa optimizes a four-dimensional objective where spatial geometry changes over time, forcing latents to encode causal physical motion rather than appearance cues. With only single-view RGB-D input, GeoMoLa attains state-of-the-art results across manipulation benchmarks and transfers to novel scenes with physically consistent transformations. Real-world tests in cluttered environments confirm robust control with minimal demonstrations, supported by ablations that identify geometric prediction as the key driver of performance.","GeoMoLa: Geometry-Aware Motion Latents for Learning Robust Manipulation  \nPolicies  \nYunchao Zhang 1 Yijia Weng 2 Ruizhe Liu 1 Ming Hu 3 Leonidas Guibas 2 Yanchao Yang 1  \narXiv :2607 .047 14v 1 [ cs .RO] 6 Jul 2026  \nAbstract  \nLearning motion latents for robotic manipulation heavily relies on extracting motion patterns from visual sequences, yet effective action abstractions require understanding three-dimensional geometric transformations. Here, we introduce GeoMoLa (Geometry-Aware Motion Latents), which learns discrete motion latent codes by predicting how point clouds evolve during manipulation rather than reconstructing visual observations. This fourdimensional objective – spatial geometry changing through time – forces latent representations to encode actual physical motion rather than appearance patterns. GeoMoLa achieves state-of-the-art performance using only single-view RGB-D input, while existing methods require multi-view reconstruction, succeeding across diverse manipulation benchmarks. Our ablations reveal that geometric prediction is the key to driving performance, quantitatively validating that manipulation depends on spatial understanding. Furthermore, the learned codes exhibit effective motion abstraction: applying them to novel scenes produces physically consistent transformations regardless of visual context. Our real-world experiments also confirm this robustness capability, achieving robust manipulation with minimal demonstrations in cluttered environments where geometric reasoning determines success. Thus, we demonstrate that effective motion latents for robot control can better emerge from understanding motion through its three-dimensional effects rather than pixel-level patterns.  \n1Department of Computer Science, The University of Hong Kong, Hong Kong, China 2Department of Computer Science, Stanford University, Stanford, California, USA 3Department of Data Science and AI, Monash University, Melbourne, Australia. Correspondence to: Yanchao Yang \u003Cycyang@cs.hku.hk> .  \nProceedings of the 43 rd International Conference on Machine Learning, Seoul, South Korea. PMLR 306, 2026 . Copyright 2026 by the author(s) .  \n1. Introduction  \nRobot manipulation requires learning reusable motion patterns – motion latents (Bruce et al., 2024 ; Parker-Holder et al., 2024 ; Ball et al., 2025)– that abstract complex continuous movements into discrete, transferable skills. Current methods mainly learn these motion latents from sequences of two-dimensional images, missing the three-dimensional geometric structure that fundamentally determines manipulation success. A grasping action, for instance, depends not only on visual appearance but also on precise spatial relationships, approach angles, and the continuous evolution of three-dimensional configurations over time.  \nTherefore, learning motion latents without access to this underlying spatiotemporal geometry may produce representations that fail to generalize across different viewpoints, object poses, or spatial arrangements. Furthermore, this representational gap could create cascading failures in realworld deployment. Robots may not recognize that occluded objects maintain their geometric relationships despite visual changes, and small spatial errors compound across action sequences without understanding of three-dimensional workspace dynamics. Solving this essential representation issue in motion latent learning is necessary for robots to understand manipulation through spatial relationships and physical transformations rather than pixel patterns.  \nThe core challenge lies in jointly modeling spatial geometry and temporal dynamics without prohibitive computational cost. Existing methods capture either spatial structure through static three-dimensional representations (Ke et al., 2024 ; Ze et al., 2024) or temporal patterns through twodimensional video (Ye et al., 2024 ; Chen et al., 2024), but not both. Three-dimensional approaches process frozen point clouds without mo","cbCaindaijgzQ3Cq","https://ap.wps.com/l/cbCaindaijgzQ3Cq","pdf",2425053,1,19,"English","en",105,"# Introduction\n## Representation Gap in Motion Latent Learning\n## Proposed Approach and Core Insight\n## Motion Latents from 3D vs 2D","[{\"question\":\"GeoMoLa learns motion latents using what training objective?\",\"answer\":\"GeoMoLa learns discrete motion latent codes by predicting how 3D point clouds evolve during manipulation, focusing on future geometric states rather than reconstructing current visual observations.\"},{\"question\":\"Why does GeoMoLa emphasize 3D geometry over 2D visual appearance?\",\"answer\":\"Robotic grasping and manipulation depend on spatial relationships and the continuous evolution of 3D configurations; without geometry, representations may fail to generalize across viewpoints, poses, and spatial arrangements.\"},{\"question\":\"How does GeoMoLa achieve strong performance and robustness with limited input?\",\"answer\":\"GeoMoLa reaches state-of-the-art results using only single-view RGB-D, and its learned codes produce physically consistent transformations in novel scenes. Real-world experiments show robustness in cluttered environments with minimal demonstrations.\"}]",1784183399,48,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"geometry-aware-motion-latents-for-learning-robust-manipulation-policies","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/geometry-aware-motion-latents-for-learning-robust-manipulation-policies/82848/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"GeoMoLa learns motion latents using what training objective?","Question",{"text":75,"@type":76},"GeoMoLa learns discrete motion latent codes by predicting how 3D point clouds evolve during manipulation, focusing on future geometric states rather than reconstructing current visual observations.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Why does GeoMoLa emphasize 3D geometry over 2D visual appearance?",{"text":80,"@type":76},"Robotic grasping and manipulation depend on spatial relationships and the continuous evolution of 3D configurations; without geometry, representations may fail to generalize across viewpoints, poses, and spatial arrangements.",{"name":82,"@type":73,"acceptedAnswer":83},"How does GeoMoLa achieve strong performance and robustness with limited input?",{"text":84,"@type":76},"GeoMoLa reaches state-of-the-art results using only single-view RGB-D, and its learned codes produce physically consistent transformations in novel scenes. Real-world experiments show robustness in cluttered environments with minimal demonstrations.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":21,"doc_module":4,"doc_module_name":45,"category_name":136,"show_sort_weight":106,"slug":137},"General","general"]