[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84482-en":3,"doc-seo-84482-105":28,"detail-sidebar-cat-0-en-105":89},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":11,"language":21,"language_code":22,"site_id":23,"html_lang":22,"table_of_contents":24,"faqs":25,"seo_title":13,"seo_description":14,"update_tm":26,"read_time":27},84482,687197100911,"Himbo","https://ap-avatar.wpscdn.com/avatar/a000239b6f1da00475?x-image-process=image/resize,m_fixed,w_180,h_180&k=1782698725881665579",8,"Research & Report","Learning Whole-Body Humanoid Locomotion via Motion Generation and Motion Tracking","Whole-body humanoid locomotion demands high-dimensional control, morphological stability, and real-time terrain adaptation from onboard perception. Reward-shaped reinforcement learning often overemphasizes lower-body behaviors, while imitation-based RL usually replays reference motions without online, perception-driven adaptation. The proposed framework combines diffusion-based motion generation for terrain-aware references, reinforcement-trained whole-body reference tracking, and closed-loop fine-tuning using a frozen motion generator to handle imperfect references. Directional goal-reaching with terrain-aware adaptation is validated on a Unitree G1 robot across boxes, hurdles, stairs, and mixed terrain, with quantitative gains in generalization and robustness.","Learning Whole-Body Humanoid Locomotion via Motion Generation and Motion Tracking  \nZewei Zhang 1 ,4 , Kehan Wen 1 , Michael Xu2 , Junzhe He 1 , Chenhao Li 1 ,3 , Takahiro Miki 1 , Clemens Schwarke 1 ,5 , Chong Zhang 1 ,3 , Xue Bin Peng2 ,5 , and Marco Hutter1  \narXiv :2604 . 17335v2 [ cs .RO] 12 Jul 2026  \nAbstract—Whole-body humanoid locomotion is challenging due to high-dimensional control, morphological instability, and the need for real-time adaptation to various terrains using onboard perception. Directly applying reinforcement learning (RL) with reward shaping to humanoid locomotion often leads to lower-body-dominated behaviors, whereas imitation-based RL can learn more coordinated whole-body skills but is typically limited to replaying reference motions without a mechanism to adapt them online from perception for terrain-aware locomotion. To address this gap, we propose a whole-body humanoid locomotion framework that combines skills learned from reference motions with terrain-aware adaptation. We first train a diffusion model on retargeted human motions for real-time prediction of terrainaware reference motions. Concurrently, we train a whole-body reference tracker with RL using this motion data. To improve robustness under imperfectly generated references, we further fine-tune the tracker with a frozen motion generator in a closed-loop setting. The resulting system supports directional goal-reaching control with terrain-aware whole-body adaptation, and can be deployed on a Unitree G1 humanoid robot with onboard perception and computation. The hardware experiments demonstrate successful traversal over boxes, hurdles, stairs, and mixed terrain combinations. Quantitative results further show the benefits of incorporating online motion generation and fine-tuning the motion tracker for improved generalization and robustness.  \nI. INTRODUCTION  \nDeveloping whole-body perceptive humanoid locomotion remains a challenging yet important step toward expanding the operational range of humanoid robots. While deep reinforcement learning (RL) has achieved remarkable success in enabling highly dynamic behaviors on quadrupedal systems [1– 6], extending these capabilities to humanoid robots remains considerably difficult, as humanoids possess higher degrees of freedom together with more complex morphological constraints. As a result, training an RL-based whole-body humanoid controller solely through reward shaping can be non-trivial and often suffers from inefficient exploration. Without explicit structural guidance, vanilla RL-based humanoid controllers frequently converge to locomotion strategies with limited wholebody coordination, making it difficult to acquire coordinated behaviors for challenging terrain traversal. This limitation becomes especially evident in tasks such as humanoid parkour,  \nManuscript received: April 9, 2026; Accepted: June 18, 2026 . This paper was recommended for publication by Editor Olivier Stasse upon evaluation of the Associate Editor and Reviewers comments.  \n1Robotic Systems Lab, ETH Zurich {zewzhang, kehwen, junzhe, chenhli, tamiki, cschwarke, chozhang, [mahutter}@ethz.ch](mahutter}@ethz.ch)  \n2 Simon Fraser University {mxa23, [xbpeng}@sfu.ca](xbpeng}@sfu.ca)  \n3ETH AI Center. 4EPFL. 5NVIDIA.  \nDigital Object Identifier (DOI): see top of this page.  \nwhere coordinated whole-body movements are essential for traversing large and complex obstacles.  \nTo alleviate the difficulty of reward shaping and inefficient exploration, motion imitation via RL has emerged as a promising alternative [7–12] . By leveraging reference data, motion tracking can efficiently transfer highly coordinated whole-body skills to the robot. However, pure motion trackers are inherently limited to replaying choreographed trajectories and generally lack the adaptability and reactivity required for operation in unstructured and diverse environments. Achieving human-level perceptive locomotion therefore requires a higherlevel compositi","cbCairRkslwD5YCG","https://ap.wps.com/l/cbCairRkslwD5YCG","pdf",1736744,1,"English","en",105,"# Introduction\n## Challenges in whole-body humanoid locomotion\n## Limits of reward shaping and imitation-based RL\n## Skill composition and generative-model approaches\n## Proposed motion generation + motion tracking framework","[{\"question\":\"Why is reward-shaped reinforcement learning difficult for whole-body humanoid locomotion?\",\"answer\":\"It can be hard to train due to inefficient exploration and may converge to strategies dominated by the lower body, limiting coordinated whole-body behavior on complex terrain.\"},{\"question\":\"What gap does imitation-based RL leave for terrain-aware locomotion?\",\"answer\":\"It typically replays reference motions and lacks an online mechanism to adapt those motions using real-time perceptive inputs for changing terrains.\"},{\"question\":\"How does the proposed framework enable terrain-aware whole-body adaptation?\",\"answer\":\"It trains a diffusion model to predict terrain-aware reference motions in real time, trains a whole-body reference tracker with RL, and then fine-tunes the tracker in a closed loop with a frozen generator to improve robustness when references are imperfect.\"}]",1784195959,20,{"code":4,"msg":29,"data":30},"ok",{"site_id":23,"language":22,"slug":31,"title":13,"keywords":32,"description":14,"schema_data":33,"social_meta":84,"head_meta":86,"extra_data":88,"updated_unix":26},"learning-whole-body-humanoid-locomotion-via-motion-generation-and-motion-tracking","",{"@graph":34,"@context":83},[35,52,66],{"@type":36,"itemListElement":37},"BreadcrumbList",[38,42,46,49],{"item":39,"name":40,"@type":41,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":43,"name":44,"@type":41,"position":45},"https://docshare.wps.com/document/","Document",2,{"item":47,"name":12,"@type":41,"position":48},"https://docshare.wps.com/document/research-report/",3,{"item":50,"name":13,"@type":41,"position":51},"https://docshare.wps.com/document/learning-whole-body-humanoid-locomotion-via-motion-generation-and-motion-tracking/84482/",4,{"url":50,"name":13,"@type":53,"author":54,"headline":13,"publisher":56,"fileFormat":59,"inLanguage":22,"description":14,"dateModified":60,"datePublished":60,"encodingFormat":59,"isAccessibleForFree":61,"interactionStatistic":62},"DigitalDocument",{"name":9,"@type":55},"Person",{"url":39,"name":57,"@type":58},"DocShare","Organization","application/pdf","2026-07-16",true,{"@type":63,"interactionType":64,"userInteractionCount":4},"InteractionCounter",{"@type":65},"ViewAction",{"@type":67,"mainEntity":68},"FAQPage",[69,75,79],{"name":70,"@type":71,"acceptedAnswer":72},"Why is reward-shaped reinforcement learning difficult for whole-body humanoid locomotion?","Question",{"text":73,"@type":74},"It can be hard to train due to inefficient exploration and may converge to strategies dominated by the lower body, limiting coordinated whole-body behavior on complex terrain.","Answer",{"name":76,"@type":71,"acceptedAnswer":77},"What gap does imitation-based RL leave for terrain-aware locomotion?",{"text":78,"@type":74},"It typically replays reference motions and lacks an online mechanism to adapt those motions using real-time perceptive inputs for changing terrains.",{"name":80,"@type":71,"acceptedAnswer":81},"How does the proposed framework enable terrain-aware whole-body adaptation?",{"text":82,"@type":74},"It trains a diffusion model to predict terrain-aware reference motions in real time, trains a whole-body reference tracker with RL, and then fine-tunes the tracker in a closed loop with a frozen generator to improve robustness when references are imperfect.","https://schema.org",{"og:url":50,"og:type":85,"og:title":13,"og:site_name":57,"og:description":14},"article",{"robots":87,"canonical":50},"index,follow",{"doc_id":7,"site_id":23},{"code":4,"msg":5,"data":90},[91,95,99,103,108,113,118,121,125,128,132],{"id":20,"doc_module":4,"doc_module_name":44,"category_name":92,"show_sort_weight":93,"slug":94},"Story & Novel",90,"story-novel",{"id":45,"doc_module":4,"doc_module_name":44,"category_name":96,"show_sort_weight":97,"slug":98},"Literature",80,"literature",{"id":51,"doc_module":4,"doc_module_name":44,"category_name":100,"show_sort_weight":101,"slug":102},"Exam",70,"exam",{"id":104,"doc_module":4,"doc_module_name":44,"category_name":105,"show_sort_weight":106,"slug":107},5,"Comic",60,"comic",{"id":109,"doc_module":4,"doc_module_name":44,"category_name":110,"show_sort_weight":111,"slug":112},6,"Technology",50,"technology",{"id":114,"doc_module":4,"doc_module_name":44,"category_name":115,"show_sort_weight":116,"slug":117},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":44,"category_name":12,"show_sort_weight":119,"slug":120},30,"research-report",{"id":122,"doc_module":4,"doc_module_name":44,"category_name":123,"show_sort_weight":27,"slug":124},9,"Religion & Spirituality","religion-spirituality",{"id":27,"doc_module":4,"doc_module_name":44,"category_name":126,"show_sort_weight":27,"slug":127},"World Cup","world-cup",{"id":129,"doc_module":4,"doc_module_name":44,"category_name":130,"show_sort_weight":129,"slug":131},10,"Lifestyle","lifestyle",{"id":133,"doc_module":4,"doc_module_name":44,"category_name":134,"show_sort_weight":104,"slug":135},19,"General","general"]