[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-81576-en":3,"doc-seo-81576-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},81576,962075006959,"Anda","https://ap-avatar.wpscdn.com/avatar/e0002397efbe92a78e?_k=1776741047341049297",8,"Research & Report","ReinforceGen: Hybrid Skill Policies with Automated Data Generation and Reinforcement Learning","Long-horizon manipulation remains difficult in robotics due to costly demonstrations and sparse reward exploration. ReinforceGen introduces a hybrid framework that segments tasks into localized skills connected via motion planning, then trains skills and planning targets with imitation learning from a small set of human demonstrations. The system improves each component through online reinforcement-learning-based fine-tuning and adaptation. On Robosuite, ReinforceGen achieves 80% success across visuomotor controls and ablations show up to 89% average performance gains, with stronger real-world transfer after fine-tuning.","arXiv :2512 . 16861v2 [ cs .RO] 10 Jul 2026  \nReinforceGen: Hybrid Skill Policies with Automated Data Generation and Reinforcement Learning  \nZihan Zhou 1 , Animesh Garg2 , Ajay Mandlekar3∗, Caelan Garrett3 ∗ , 1 University of Toronto, Vector Institute 2 Georgia Institute of Technology 3 NVIDIA Research  \n∗ equal advising  \nAbstract: Long-horizon manipulation has been a long-standing challenge in the robotics community. We propose ReinforceGen, a system that combines task decomposition, data generation, imitation learning, and motion planning to forman initial solution, and improves each component through reinforcement-learningbased fine-tuning. ReinforceGen first segments the task into multiple localized skills, which are connected through motion planning. The skills and motion planning targets are trained with imitation learning on a dataset generated from 10 human demonstrations, and then fine-tuned through online adaptation and reinforcement learning. When benchmarked on the Robosuite dataset, ReinforceGen reaches 80% success rate on all tasks with visuomotor controls in the highest reset range setting. Additional ablation studies show that our fine-tuning approaches contribute to an 89% average performance increase. Finally, ReinforceGen demonstrates significant improvement through fine-tuning in our real-world evaluations.  \nMore results and videos are available [at](at reinforcegen.github.io)[ reinforcegen.github.io](at reinforcegen.github.io).  \nKeywords: Reinforcement Learning, Manipulation Planning, Data Generation  \n1 Introduction  \nImitation Learning (IL) from demonstrations is an effective approach for robots to autonomously complete tasks. In long-horizon tasks, collecting demonstrations can be expensive, and the trained agent is more likely to deviate from the demonstrations to out-of-distribution states. Reinforcement Learning (RL) leverages random exploration, incorporating environmental feedback through rewards. However, long horizons exacerbate the exploration challenge, especially when the reward signals are also sparse. In the context of robot learning, collecting demonstrations for IL is often time-consuming and expensive, as it typically requires a teleoperation platform and coordinating with human operators. Furthermore, the solution quality and data coverage of the demonstrations are critical, as they directly impact the agent’s performance when used in methods such as Behavior Cloning (BC) [1] .  \nOne approach to combat demonstration insufficiency is to augment the dataset through synthetic data generation. In robotic manipulation tasks, a line of work [2–4] focuses on object-centric data generation through demonstration adaptation. Other approaches [5, 6] use Task and Motion Planning (TAMP) [7] to generate demonstrations. An alternative strategy is to hierarchically divide the task into consecutive stages with easier-to-solve subgoals [3, 8, 9] . In most manipulation tasks, only a small fraction of robot execution requires high-precision movements, for example, only the contact-rich segments. Thus, these approaches concentrate the demo collection effort at the precision-intensive skill segments and connect segments using planning, ultimately improving demo sample efficiency.  \nStill, these demonstration generation methods are open-loop and rely solely on offline data. As a result, IL agents trained with the generated data are still bottlenecked by the quality of the source demonstrations. To combat this, we propose ReinforceGen, a framework that improves hierarchical data generation by incorporating online exploration and environmental feedback using RL. ReinforceGen trains a hybrid BC agent with object-centric data generation as its base policy. It then combines distillation, causal inference, and RL to improve the base agent with online data, as well as real-time adaptation from environment feedback during deployment. We demonstrate that ReinforceGen produces high-performance hybrid data generators","cbCaiuWnbFqFvC0U","https://ap.wps.com/l/cbCaiuWnbFqFvC0U","pdf",7626382,2,1,21,"English","en",105,"# Introduction\n# Related Work","[{\"question\":\"What problem does ReinforceGen address in robotics manipulation?\",\"answer\":\"ReinforceGen targets long-horizon manipulation where collecting demonstrations is expensive and reinforcement learning suffers from sparse-reward exploration.\"},{\"question\":\"How does ReinforceGen generate and improve training data?\",\"answer\":\"It first segments tasks into localized skills, generates an offline dataset via synthetic data from a small set of human demonstrations, trains with imitation learning, and then improves via online adaptation and reinforcement learning.\"},{\"question\":\"What performance results does ReinforceGen report?\",\"answer\":\"On the Robosuite benchmark, ReinforceGen reaches an 80% success rate on all tasks with visuomotor controls under the highest reset range setting, and fine-tuning yields an 89% average performance increase in ablation studies.\"}]",1784174416,53,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"reinforcegen-hybrid-skill-policies-with-automated-data-generation-and-reinforcement-learning","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/reinforcegen-hybrid-skill-policies-with-automated-data-generation-and-reinforcement-learning/81576/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does ReinforceGen address in robotics manipulation?","Question",{"text":75,"@type":76},"ReinforceGen targets long-horizon manipulation where collecting demonstrations is expensive and reinforcement learning suffers from sparse-reward exploration.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does ReinforceGen generate and improve training data?",{"text":80,"@type":76},"It first segments tasks into localized skills, generates an offline dataset via synthetic data from a small set of human demonstrations, trains with imitation learning, and then improves via online adaptation and reinforcement learning.",{"name":82,"@type":73,"acceptedAnswer":83},"What performance results does ReinforceGen report?",{"text":84,"@type":76},"On the Robosuite benchmark, ReinforceGen reaches an 80% success rate on all tasks with visuomotor controls under the highest reset range setting, and fine-tuning yields an 89% average performance increase in ablation studies.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]