[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83864-en":3,"doc-seo-83864-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83864,8796095462418,"Noah","https://ap-avatar.wpscdn.com/avatar/80000253c1241d02b47?x-image-process=image/resize,m_fixed,w_180,h_180&k=1778826106357471780",8,"Research & Report","Athena-WBC: Capability-Aligned Policy Experts for Long-Tail Humanoid Whole-Body Control","Large-scale humanoid motion-tracking controllers can be improved by resampling and specializing hard motions, but that perspective misses a key failure pattern: even strong whole-body-control baselines leave a residual set of feasible training clips unsolved. These failures concentrate in high-dynamic transitions and balance-critical behaviors, stemming from a capability mismatch induced by the default training recipe. Athena-WBC introduces capability-aligned dynamic and balance expert teachers, distills them into a single deployable controller, and further refines it with RL. Experiments on a full-size humanoid show better recovery and held-out tracking than a strong SONIC-recipe baseline using only a small number of experts.","arXiv :2607 .04837v2 [ cs .RO] 7 Jul 2026  \nXPENG ROBOTICS  \nAthena-WBC: Capability-Aligned Policy Experts for LongTail Humanoid Whole-Body Control  \nYuan Jiang∗ , Ningyuan Zhang∗ , Xicun Yang, Yuzhi Jiang, Jie Chen†  \nXPENG Robotics  \nAbstract  \nLarge-scale humanoid motion-tracking controllers are commonly improved by reallocating training effort: difficult motions are sampled more often, isolated into smaller subsets, or assigned to specialized experts. We show that this view is incomplete. In strong whole-body-control baselines, a residual set of feasible training clips remains unsolved even under targeted training, especially for high-dynamic transitions and balance-critical motions. These failures arise not only from insufficient exposure, but from a mismatch between the motion demands and the effective capability induced by the default training recipe.  \nWe propose Athena-WBC, a compact teacher-student pipeline with capability-aligned policy experts for long-tail humanoid whole-body control. Dynamic experts use a tracking-focused, constraintaware objective that removes conservative effort and temporal-control penalties while preserving physical feasibility constraints; balance experts use a gravity curriculum to improve early-training survivability. The resulting privileged teachers are motion-routed for DAgger distillation and then compressed into a single controller with deployable observations followed by RL fine-tuning. Experiments on a full-size humanoid show improved recovery of training-set long-tail motions and better held-out tracking than a strong SONIC-recipe baseline, using only a small number of experts.  \n1 Introduction  \nHumanoid whole-body control (WBC) has progressed from tracking isolated skills to imitating large and diverse human motion corpora [1–3] . Modern motion-conditioned policies can achieve high aggregate tracking performance by combining reinforcement learning, motion tracking rewards, privileged teacher policies, and deployable student distillation. However, high average performance does not imply that the training distribution has been fully absorbed. In particular, the high-coverage regime exposes a residual training-set long tail: motions that appear in the training corpus but remain unsolved by the learned controller. This failure mode is distinct from generalization failure and has received comparatively less attention.  \nWe observe this phenomenon in the SONIC baseline. As shown in Fig. 2, evaluating the released SONIC checkpoint on its training motion set leaves a small but meaningful fraction of clips unsuccessful. These failures are not uniformly distributed. They concentrate in high-dynamic transitions, such as rapid direction changes and aggressive contact switches, and in balance-critical motions, such as low-support poses and slow recovery phases. Moreover, in our own controlled experiments, we find that a subset of failures persists even when the policy is trained only on the failed clips. This suggests that the default training recipe can induce an effective capability bottleneck rather than merely suffering from insufficient exposure to rare motions.  \nA common response to long-tail failures is to reallocate training effort: sample difficult motions more frequently [1, 4–9], cluster motions into narrower subsets [10], or train larger banks of specialized  \n* Equal contribution: Yuan Jiang and Ningyuan Zhang.  \n† Correspondence: [Jie Chen at chenj81@xiaopeng.com](Jie Chen at chenj81@xiaopeng.com)  \nFigure 1: Overview of Athena-WBC. A general privileged teacher is trained on the full motion set. Residual failures are mined and used to train dynamic and balance experts in parallel. The frozen teachers are then routed per motion, distilled into a single student, and finetuned with RL.  \nexperts [11] . These strategies can improve coverage, but they primarily change which data a policy sees. They do not necessarily change the control regime that the policy is encouraged to acquire","cbCaidPnKIQsYRVP","https://ap.wps.com/l/cbCaidPnKIQsYRVP","pdf",30762308,5,1,27,"English","en",105,"# Introduction\n## Long-tail failures in high-coverage training\n## Motivation: capability mismatch beyond data reallocation\n## Athena-WBC overview and training pipeline\n## Evaluation protocol and metrics","[{\"question\":\"What problem does the paper identify in humanoid whole-body control training?\",\"answer\":\"It identifies a residual long-tail of training motions that remain unsolved even when using high-coverage training regimes. These failures differ from generic generalization failures and persist under targeted training on the failed clips.\"},{\"question\":\"Why do residual failures persist even when training focuses only on the failed motions?\",\"answer\":\"The paper argues the default training recipe induces an effective capability bottleneck, creating a mismatch between motion demands and what the learned controller can control. This is not just a matter of insufficient exposure to rare motions.\"},{\"question\":\"How does Athena-WBC address long-tail failures?\",\"answer\":\"Athena-WBC trains capability-aligned dynamic experts with a tracking-focused, constraint-aware objective and balance experts using a gravity curriculum for early survivability. Privileged teachers are motion-routed, distilled into a single deployable controller via DAgger-style supervision, and then refined with RL.\"}]",1784191067,68,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"athena-wbc-capability-aligned-policy-experts-for-long-tail-humanoid-whole-body-control","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/athena-wbc-capability-aligned-policy-experts-for-long-tail-humanoid-whole-body-control/83864/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-26","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What problem does the paper identify in humanoid whole-body control training?","Question",{"text":76,"@type":77},"It identifies a residual long-tail of training motions that remain unsolved even when using high-coverage training regimes. These failures differ from generic generalization failures and persist under targeted training on the failed clips.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"Why do residual failures persist even when training focuses only on the failed motions?",{"text":81,"@type":77},"The paper argues the default training recipe induces an effective capability bottleneck, creating a mismatch between motion demands and what the learned controller can control. This is not just a matter of insufficient exposure to rare motions.",{"name":83,"@type":74,"acceptedAnswer":84},"How does Athena-WBC address long-tail failures?",{"text":85,"@type":77},"Athena-WBC trains capability-aligned dynamic experts with a tracking-focused, constraint-aware objective and balance experts using a gravity curriculum for early survivability. Privileged teachers are motion-routed, distilled into a single deployable controller via DAgger-style supervision, and then refined with RL.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":20,"slug":138},19,"General","general"]