[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83120-en":3,"doc-seo-83120-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83120,1099514067438,"River Wang","https://ap-avatar.wpscdn.com/avatar/100002539ee87300030?x-image-process=image/resize,m_fixed,w_180,h_180&k=1780474512215547542",8,"Research & Report","RynnWorld-Teleop: An Action-Conditioned World Model for Digital Teleoperation","RynnWorld-Teleop introduces digital teleoperation to scale robot learning when physical teleoperation limits data collection by tying each demonstration to operator time, specific hardware, and fixed workspaces. The method replaces the real robot with a generative, robot-centric world model: an operator hand-pose stream conditions synthesis of high-fidelity egocentric videos from a single reference image. The recorded pose stream becomes an embodiment-agnostic action label, enabling retargeting to diverse target robots and producing complete state-action trajectories for imitation learning. RynnWorld-Teleop combines depth-aware skeletal conditioning, video Diffusion Transformer training, and streaming autoregressive distillation, achieving 40+ FPS on one H100. Policies trained on its generated data support effective zero-shot Sim2Real transfer and augmenting real data consistently improves success rates.","arXiv :2607 .06558v2 [ cs .RO] 12 Jul 2026  \nRynnWorld-Teleop: An Action-Conditioned World Model for Digital Teleoperation  \nHaoyu Zhao∗1 ,2 ,3 , Xingyue Zhao∗1 , Hangyu Li5 , Biao Gong6 , Kehan Li 1 ,4 , Siteng Huang†1 ,4 , Xin Li 1 ,4 , Deli Zhao†1 , Zhongyu Li†2 ,3  \n1 DAMO Academy, Alibaba Group, 2 Hong Kong Embodied AI Lab, 3 CUHK, 4 Hupan Lab,  \n5 Alibaba Group, 6 Ant Group  \n∗ Equal contribution, †Corresponding author  \nScaling robot learning requires massive, diverse trajectory data, yet collection is currently bottlenecked by physical teleoperation, where every demonstration binds operator time to specific hardware and workspaces. We introduce digital teleoperation, a paradigm that decouples data collection from physical constraints by replacing the real robot with a generative world model. In this framework, an operator’shand-pose stream drives a robot-centric generative world model to synthesize high-fidelity egocentric videos from a single reference image. The recorded pose stream serves as an embodiment-agnostic action label transferable to any target robot via standard retargeting, yielding complete state-action trajectories for imitation learning independent of physical hardware. We instantiate this paradigm in RynnWorld-Teleop, a system that integrates depth-aware skeletal conditioning, progressive human-torobot training on a video Diffusion Transformer (DiT), and streaming autoregressive distillation. This pipeline compresses the generative process into a single-pass inference, enabling 40+ FPS, real-time interactive generation on a single H100 GPU. Policies trained exclusively on RynnWorld-Teleopgenerated data achieve effective zero-shot Sim2Real transfer across dexterous and diverse bimanual tasks. Moreover, augmenting real-world datasets with our digitally teleoperated data consistently improves success rates, demonstrating that RynnWorld-Teleop serves as a high-fidelity, scalable data engine for the next generation of robotic agents.  \n [https://alibaba-damo-academy.github.io/RynnWorld-Teleop.github.io](https://alibaba-damo-academy.github.io/RynnWorld-Teleop.github.io)  \n [https://github.com/alibaba-damo-academy/RynnWorld-Teleop](https://github.com/alibaba-damo-academy/RynnWorld-Teleop)  \n [https://huggingface.co/Alibaba-DAMO-Academy/RynnWorld-Teleop](https://huggingface.co/Alibaba-DAMO-Academy/RynnWorld-Teleop)  \n [https://www.modelscope.cn/models/DAMO_Academy/RynnWorld-Teleop](https://www.modelscope.cn/models/DAMO_Academy/RynnWorld-Teleop)  \nDate: July 14, 2026   \n1 Introduction  \nRecent advances in robotics, such as vision-language-action (VLA) models (Black et al. , 2024 ; Brohan et al. , 2022 ; Kim et al. , 2024 ; Li et al. , 2026b ; Zhao et al. , 2026b) and world models (Agarwal et al. , 2025 ; Ali et al. , 2025), show emerging promise for general-purpose autonomy, yet they remain significantly hindered by data scarcity (Zhao et al. , 2026a ; Lepert et al. , 2025a ; Zhao et al. , 2023) . Access to such large-scale robot data would enable more straightforward training and potentially unlock a higher performance upper bound. While traditional teleoperation systems (Li et al. , 2025b ; Ze et al. , 2025 ; Zhao et al. , 2025a) provide high-quality expert data, they are often confined to fixed laboratory settings and specific objects Zhao et al.(2025b); Guo et al. (2026); Zhao et al. (2024) . The immense overhead of manual environment resets and the logistical challenge of procuring diverse real-world objects prevent these systems from capturing the long-tail distribution of interactions. Consequently, achieving the robust generalization required for unstructured environments remains an open challenge that physical platforms alone cannot solve.  \nWe propose to decouple operator time from physical infrastructure by replacing the real robot with a generative one. We call this paradigm digital teleoperation: an operator’s real-time hand-pose stream is consumed by  \nPhysical Teleoperation  \nRetargeting Control","cbCaitO3hHhKLVnq","https://ap.wps.com/l/cbCaitO3hHhKLVnq","pdf",12961004,3,1,23,"English","en",105,"# Introduction\n## Digital teleoperation vs. physical teleoperation\n## Core paradigm and data generation","[{\"question\":\"What problem does digital teleoperation address compared with physical teleoperation?\",\"answer\":\"Physical teleoperation binds each demonstration to operator time, specific hardware, and fixed workspaces, limiting throughput. Digital teleoperation decouples data collection from physical constraints by using a generative world model instead of a real robot.\"},{\"question\":\"How does RynnWorld-Teleop generate training data without moving a real robot?\",\"answer\":\"An operator’s real-time hand-pose stream conditions a robot-centric generative world model. With a single reference image, the model synthesizes the egocentric video the robot would have produced, yielding aligned state-action trajectories for imitation learning.\"},{\"question\":\"How can the same gesture be used across different robot embodiments?\",\"answer\":\"The recorded hand-pose stream is treated as an embodiment-agnostic, ground-truth action label. Standard retargeting transfers this label into embodiment-specific robot actions, enabling flexible reuse of collected behaviors.\"}]",1784185412,58,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"rynnworld-teleop-an-action-conditioned-world-model-for-digital-teleoperation","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/rynnworld-teleop-an-action-conditioned-world-model-for-digital-teleoperation/83120/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does digital teleoperation address compared with physical teleoperation?","Question",{"text":75,"@type":76},"Physical teleoperation binds each demonstration to operator time, specific hardware, and fixed workspaces, limiting throughput. Digital teleoperation decouples data collection from physical constraints by using a generative world model instead of a real robot.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does RynnWorld-Teleop generate training data without moving a real robot?",{"text":80,"@type":76},"An operator’s real-time hand-pose stream conditions a robot-centric generative world model. With a single reference image, the model synthesizes the egocentric video the robot would have produced, yielding aligned state-action trajectories for imitation learning.",{"name":82,"@type":73,"acceptedAnswer":83},"How can the same gesture be used across different robot embodiments?",{"text":84,"@type":76},"The recorded hand-pose stream is treated as an embodiment-agnostic, ground-truth action label. Standard retargeting transfers this label into embodiment-specific robot actions, enabling flexible reuse of collected behaviors.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]