[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83427-en":3,"doc-seo-83427-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83427,7971461741311,"Ophelia","https://ap-avatar.wpscdn.com/avatar/74000253aff267980c6?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779345379180704826",8,"Research & Report","ARDY Autoregressive Diffusion with Hybrid Representation for Interactive Human Motion Generation","ARDY introduces an autoregressive diffusion model for interactive 3D human motion generation that supports online text prompting together with flexible kinematic constraints across long horizons. The method natively handles root waypoints and trajectories, full-body keyframes, and sparse joint positions and rotations, producing controllable and responsive motion from real-time mouse and keyboard inputs. A 4-step diffusion design achieves an average 33 ms generation latency, improving inference speed while preserving motion realism, constraint adherence, and responsiveness in interactive applications.","arXiv :2607 .0874 1v 1 [ cs .GR] 9 Jul 2026  \nARDY: Autoregressive Diffusion with Hybrid Representation for Interactive Human Motion Generation  \nKAIFENG ZHAO, NVIDIA, Switzerland and ETH Zürich, Switzerland MATHIS PETROVICH, NVIDIA, Switzerland  \nHAOTIAN ZHANG, NVIDIA, USA TINGWU WANG, NVIDIA, USA SIYU TANG, ETH Zürich, Switzerland DAVIS REMPE, NVIDIA, USA  \nRoot velocity control via keyboard  \nEnd-effector joints  \nRoot  \ntrajectory  \nFull body  \nRoot keyframe  \nwaypoint  \n| A happy person, leaning to the left, jogs ina fast circular arc to their left, then stop. | A person is stealthily walking sideways to their left at a slow pace. | A person opens a door and walks out the door before closing it behind. | \u003Cbr>……\u003Cbr> |\n| --- | --- | --- | --- |\n\n33ms  \nFig. 1. We present ARDY, an autoregressive diffusion model designed for interactive human motion generation. Our approach natively supports online text prompting alongside a comprehensive suite of flexible kinematic constraints — including root waypoints and trajectories, full-body keyframes, and sparse joint positions and rotations — over long horizons. ARDY enables controllable and responsive interactive motion synthesis from real-time user inputs such as mouse and keyboard commands, with our efficient 4-step diffusion model achieving an average generation latency of 33 ms.  \nGenerating realistic 3D human motions in real-time within interactive applications is key for animation, simulation, and humanoid robotics. While recent offline motion generation approaches offer precise control via text and kinematic constraints, they lack the inference speed required for interactive settings. Conversely, existing online methods enable real-time synthesis but often sacrifice controllability or struggle with complex text semantics and long-horizon goals due to limited context windows. In this work, we introduce ARDY, a streaming generation framework that bridges this gap by enabling high-fidelity motion generation controllable via online text prompts and flexible kinematic constraints. ARDY employs a hybrid representation that combines explicit root features with a latent body embedding, balancing precise trajectory control with efficient generative learning. We  \nAuthors’ Contact Information: Kaifeng Zhao, [kaifeng.zhao@inf.ethz.ch](kaifeng.zhao@inf.ethz.ch), NVIDIA, Switzerland and ETH Zürich, Switzerland; Mathis Petrovich, [mpetrovich@nvidia.com](mpetrovich@nvidia.com), NVIDIA, Switzerland; Haotian Zhang, [haotianz@nvidia.com](haotianz@nvidia.com), NVIDIA, USA; Tingwu  \nWang, [tingwuw@nvidia.com](tingwuw@nvidia.com), NVIDIA, USA; Siyu Tang, [siyu.tang@inf.ethz.ch](siyu.tang@inf.ethz.ch), ETH  \nZürich, Switzerland; Davis Rempe, [drempe@nvidia.com](drempe@nvidia.com), NVIDIA, USA.  \nThis work is licensed under a Creative Commons Attribution 4 .0 International License.© 2026 Copyright held by the owner/author(s) .  \nACM 1557-7368/2026/7-ART86  \n[https://doi.org/10.1145/3811284](https://doi.org/10.1145/3811284)  \npropose a two-stage autoregressive transformer denoiser that features variable history context and supports conditioning on flexible, long-horizon kinematic constraints. By training on a large-scale motion capture dataset and being directly conditioned on text labels and kinematic constraints sampled from ground truth poses, ARDY natively learns controllable generation that supports online prompting and flexible long-horizon goals. Extensive evaluations on the HumanML3D benchmark and the large-scale, high-fidelity Bones Rigplay dataset demonstrate ARDY’s high motion quality and constraint adherence, validating the efficacy of our key architectural decisions. Finally, we demonstrate the method’s practical versatility through an interactive demo featuring dynamic text control, diverse keyframe pose constraints, path following, and interactive locomotion control via mouse and keyboard. Supplementary video results, code, and model releases can be found at [https://research.nvidia.c","cbCaiup9d4SaMSBh","https://ap.wps.com/l/cbCaiup9d4SaMSBh","pdf",7421876,3,1,14,"English","en",105,"# Introduction\n## Interactive motion generation requirements\n## Offline versus online motion modeling\n## ARDY approach overview","[{\"question\":\"What is ARDY and what problem does it address?\",\"answer\":\"ARDY is an autoregressive diffusion framework for interactive human motion generation. It bridges the gap between high-control offline methods and real-time online methods by enabling online text prompting plus flexible long-horizon kinematic constraints.\"},{\"question\":\"How does ARDY allow user control during interaction?\",\"answer\":\"ARDY supports real-time user inputs such as mouse and keyboard commands and enables online text prompting. It also conditions generation on kinematic constraints like root waypoints/trajectories, full-body keyframes, and sparse joint positions and rotations.\"},{\"question\":\"What performance and validation results are reported?\",\"answer\":\"An efficient 4-step diffusion model targets an average generation latency of about 33 ms. Evaluations on HumanML3D and the Bones Rigplay dataset show high motion quality and strong constraint adherence, and an interactive demo demonstrates dynamic text control and interactive locomotion control.\"}]",1784187523,35,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"ardy-autoregressive-diffusion-with-hybrid-representation-for-interactive-human-motion-generation","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/ardy-autoregressive-diffusion-with-hybrid-representation-for-interactive-human-motion-generation/83427/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is ARDY and what problem does it address?","Question",{"text":75,"@type":76},"ARDY is an autoregressive diffusion framework for interactive human motion generation. It bridges the gap between high-control offline methods and real-time online methods by enabling online text prompting plus flexible long-horizon kinematic constraints.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does ARDY allow user control during interaction?",{"text":80,"@type":76},"ARDY supports real-time user inputs such as mouse and keyboard commands and enables online text prompting. It also conditions generation on kinematic constraints like root waypoints/trajectories, full-body keyframes, and sparse joint positions and rotations.",{"name":82,"@type":73,"acceptedAnswer":83},"What performance and validation results are reported?",{"text":84,"@type":76},"An efficient 4-step diffusion model targets an average generation latency of about 33 ms. Evaluations on HumanML3D and the Bones Rigplay dataset show high motion quality and strong constraint adherence, and an interactive demo demonstrates dynamic text control and interactive locomotion control.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]