[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86518-en":3,"doc-seo-86518-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":11,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},86518,687197207057,"Sage","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","A Single Diffusion-Policy Controller for Multi-Task Block Pushing with Zero-Shot Sim-to-Real Transfer","Diffusion policies provide an effective way to learn complex robotic maneuvers from behavior cloning, yet extending them to multi-task manipulation and sim-to-real transfer remains challenging. This work trains a single diffusion policy from scratch using reinforcement learning for multi-task block pushing with varying block shapes. A simplified policy loss reweights the BC-style evidence lower bound and becomes an RL learning module. Exploration without demonstrations is addressed via reverse curriculum generation and objective-centric representations, with evaluations testing zero-shot transfer across goals, shapes, weights, and surface friction.","A Single Diffusion-Policy Controller for Multi-Task Block Pushing with Zero-Shot Sim-to-Real Transfer  \nHaitong Ma 1 , Haldun Balim 1 , Yang Hu 1 , Bo Dai2 and Na Li 1  \narXiv :2607 . 10892v1 [ cs .RO] 12 Jul 2026  \nAbstract—Diffusion policies have shown promising empirical performance in representing and learning complex maneuvers for robots using behavior cloning (BC). In this paper, we explore training diffusion policies from scratch using reinforcement learning (RL) for multi-task robotic manipulation. Specifically, we aim to train a single diffusion policy for block-pushing tasks with multiple shapes. The proposed framework features a simple policy loss function, which is a reweighted evidence lower bound used in BC-based diffusion policy training and can seamlessly serve as the policy learning module in RL algorithms. To address the exploration challenges arising from the absence of demonstrations, we incorporate reverse curriculum generation and objective-centric representations. Combined with the expressiveness of diffusion policies, our design supports learning of multi-task block-pushing policies in our sparsereward simulation setting. We further evaluate whether the trained diffusion policy transfers in zero-shot to real-world tasks under varying environmental conditions including goal positions, block shapes, block weights and surface friction, providing evidence that this pipeline can transfer to our realworld block-pushing setup under the tested variations.  \nI. INTRODUCTION  \nDenoising diffusion probabilistic model (DDPM) [1], [2],[3] has been showing remarkable expressiveness and flexibility in representing and generating complex data distributions [4] . In light of these strengths, diffusion policies have been widely leveraged to generate long-horizon robot trajectories [5], [6], [7] or imitate expert policies in MDPs [8],[9], [10], [11] using behavior cloning or offline RL methods when expert demonstrations are available in abundance.  \nBuilding on this foundation, a natural question to ask is whether such expressive policy classes can go beyond singletask learning and serve as a basis for more general-purpose robotic intelligence, since the ultimate goal of autonomous robotic manipulation is to learn one single universal policy that can solve a variety of different tasks [7] . These tasks are often contact-rich, where the presence of frequent and complex contacts makes the planning landscape highly sensitive to differences in both task objectives and external disturbances. Capturing such variability requires sufficiently expressive models to parameterize the universal policies—this is precisely where diffusion policies come into play, owing to their inherent expressiveness and flexibility [12], [13] .  \nDespite the great success of diffusion policies, the current training approaches are generally based on behavior cloning (BC) [6], while RL only serves as an optional fine-tuning method [5] . However, such BC-based approaches could be challenging for multi-task manipulations, since the task space  \n1Harvard University.  \n2 Georgia Institute of Technology.  \nrequires a lot of diverse expert demonstrations across different tasks, which limits the scalability of the BC approaches. Worse still, the fine-tuning phase is usually fragile and faces the challenge of catastrophic forgetting [?], hindering its ability to reliably adapt to new tasks without compromising the performance of previously learned skills.  \nRecently, a growing interest has emerged in training diffusion policies from scratch using online RL methods [14],[15], [16], which achieves stronger performance compared to previous algorithms based on deterministic or Gaussian policies [17], [18] . In fact, online RL has shown its strong potential for multi-task learning in arcade games and robotics locomotion [19], [20] . However, multi-task diffusion has been studied only with BC-based training [12], [21] but not yet with online RL. Meanwhile, one of the","cbCailHClD5AqdpH","https://ap.wps.com/l/cbCailHClD5AqdpH","pdf",4153011,5,1,"English","en",105,"# Introduction\n## Motivation: from single-task to universal control\n## Limits of BC-based diffusion training\n## Online RL for diffusion policies and sim-to-real challenge\n## Contribution and pipeline overview","[{\"question\":\"What is the main goal of the proposed method?\",\"answer\":\"Train a single diffusion-policy controller for multi-task, contact-rich block pushing and enable zero-shot transfer from simulation to real-world settings under changing conditions.\"},{\"question\":\"How does the paper address the challenge of training without demonstrations?\",\"answer\":\"It incorporates reverse curriculum generation and objective-centric representations to improve exploration when demonstration data is absent.\"},{\"question\":\"What factors are used to evaluate zero-shot sim-to-real transfer?\",\"answer\":\"The evaluations vary goal positions, block shapes, block weights, and surface friction to test whether the trained policy transfers to real-world block pushing.\"}]",1784212332,20,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"a-single-diffusion-policy-controller-for-multi-task-block-pushing-with-zero-shot-sim-to-real-transfer","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/a-single-diffusion-policy-controller-for-multi-task-block-pushing-with-zero-shot-sim-to-real-transfer/86518/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-28","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is the main goal of the proposed method?","Question",{"text":75,"@type":76},"Train a single diffusion-policy controller for multi-task, contact-rich block pushing and enable zero-shot transfer from simulation to real-world settings under changing conditions.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the paper address the challenge of training without demonstrations?",{"text":80,"@type":76},"It incorporates reverse curriculum generation and objective-centric representations to improve exploration when demonstration data is absent.",{"name":82,"@type":73,"acceptedAnswer":83},"What factors are used to evaluate zero-shot sim-to-real transfer?",{"text":84,"@type":76},"The evaluations vary goal positions, block shapes, block weights, and surface friction to test whether the trained policy transfers to real-world block pushing.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,109,114,119,122,126,129,133],{"id":21,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":45,"category_name":106,"show_sort_weight":107,"slug":108},"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":45,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":28,"slug":125},9,"Religion & Spirituality","religion-spirituality",{"id":28,"doc_module":4,"doc_module_name":45,"category_name":127,"show_sort_weight":28,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":45,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":45,"category_name":135,"show_sort_weight":20,"slug":136},19,"General","general"]