[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84187-en":3,"doc-seo-84187-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":11,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},84187,1374391974468,"Eden","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","Compositional Motion Generation from Demonstration with Object-Centric Neural Fields","Compositionality enables scalable, data-efficient robot learning by combining simpler elements into complex behaviors. The framework presents generative learning from demonstration that links perception and motion via shared object-level representations. Scenes are rendered using object-centric neural fields combining canonical neural fields with latent-conditioned deformations, yielding smooth, consistent, interpretable spatial and geometric variations. Motion generation uses a temporal mixture-of-experts gating mechanism to combine object-conditioned movement primitives into full trajectories. Simulation completes long-horizon manipulation with less training data, while real-world tests show robustness to noise, category-level generalization with language segmentation, and operation directly on 3D scene representations.","Compositional Motion Generation from Demonstration with Object-Centric Neural Fields  \nAhmet Tekden 1 Yasemin Bekiroglu 1 ,2  \narXiv :2607 .07 129v 1 [ cs .RO] 8 Jul 2026  \nAbstract—Compositionality, by organizing complex behavior as combinations of simpler elements, enables robot learning that is scalable and data efficient. Leveraging this principle, we propose a generative learning-from-demonstration framework that enables compositional modeling of robotic behavior by connecting perception and motion through shared object-level representations. We render scenes from object-centric neural representations that integrate canonical neural fields with latentconditioned deformations, capturing positional and geometric variations in a smooth, consistent, and interpretable way. For motion generation, a temporal mixture-of-experts (MoE) employs a gating mechanism to combine object-conditioned movement primitives over time, producing complete trajectories. This spatial–temporal compositionality maintains the data efficiency of movement primitives while grounding motion in visual structure, enabling systematic generalization across diverse scene configurations. In simulation, long-horizon manipulation tasks are successfully completed using the proposed model, which requires significantly less training data than other image-based baselines. Real-world experiments further demonstrate the method’s robustness to noise, its ability to generalize at the category level through language-based segmentation models, and its capacity to operate directly on 3D scene representations.  \nIndex Terms—Learning from Demonstration, Deep Learning in Grasping and Manipulation  \nI. INTRODUCTION  \nLearning from Demonstration (LfD) [1] is a widely used  \napproach for teaching robots task-specific skills from human demonstrations. A central challenge is representing diverse, long-horizon behaviors from a limited number of demonstrations. Treating motion as a single, monolithic sequence overlooks its compositional structure, which complicates the learning process. Everyday tasks unfold through sequential, object-level interactions—reaching, grasping, moving, and placing—that together accomplish the overall goal. This motion-level compositionality [2] can be complemented by scene-level compositionality, grounding motion in objectcentric scene representations and providing a natural foundation for compositional motion modeling. Figure 1 illustratesan example of such object-centric structure, showing how the scene is decomposed into object-specific representationsand how those representations relate to different phases of a  \nManuscript received: December 18, 2025; Revised May 25, 2026; Accepted June 24, 2026 .  \nThis paper was recommended for publication by Editor Wei Pan upon evaluation of the Associate Editor and Reviewers’ comments.  \nThis work was partially supported by the Wallenberg AI, Autonomous Systems and Software Program (WASP) funded by the Knut and Alice Wallenberg Foundation, the Chalmers AI Research Center (CHAIR), and the Chalmers Gender Initiative for Excellence (Genie) . 1Department of Electrical Engineering, Chalmers University of Technology, SE-412 96 Gothenburg, Sweden. 2Department of Computer Science, University College London,  \nWC1E 6BT London, U.K. Email: [tekden@chalmers.se](tekden@chalmers.se)[ ](tekden@chalmers.se)Digital Object Identifier (DOI): see top of this page.  \nFig. 1: System overview. Top: Task sequence with object interactions, followed by the ground-truth and reconstructed images. Middle: Object-specific masks inferred from latent codes, illustrating object-centric scene decomposition. Bottom: Temporal gating weights showing each object’s influence over the trajectory; snapshots above the plot correspond to key events (T0–T3), marked by vertical dotted lines.  \nmanipulation task. By supporting data efficiency, modularity, and generalization, compositional formulations present a compelling alternative to monolithic approache","cbCaivAXiExCIMpR","https://ap.wps.com/l/cbCaivAXiExCIMpR","pdf",3915710,6,1,"English","en",105,"# Introduction\n## Learning from Demonstration (LfD)\n## Movement Primitives and Limitations\n## Proposed Generative Object-Centric MoE Framework","[{\"question\":\"What problem does the paper address in learning from demonstrations?\",\"answer\":\"It tackles the challenge of representing diverse, long-horizon robotic behaviors from limited demonstrations, where treating motion as a single monolithic sequence makes learning difficult.\"},{\"question\":\"How does the proposed method connect perception with motion?\",\"answer\":\"It uses object-centric shared representations: scenes are rendered from object-centric neural fields, and motion generation is grounded by object-conditioned primitives combined over time.\"},{\"question\":\"How does the temporal mixture-of-experts model produce full motion trajectories?\",\"answer\":\"A gating mechanism selects and blends multiple object-conditioned movement experts over time, composing primitives into complete trajectories while maintaining data efficiency.\"}]",1784193787,20,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"compositional-motion-generation-from-demonstration-with-object-centric-neural-fields","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/compositional-motion-generation-from-demonstration-with-object-centric-neural-fields/84187/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the paper address in learning from demonstrations?","Question",{"text":75,"@type":76},"It tackles the challenge of representing diverse, long-horizon robotic behaviors from limited demonstrations, where treating motion as a single monolithic sequence makes learning difficult.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the proposed method connect perception with motion?",{"text":80,"@type":76},"It uses object-centric shared representations: scenes are rendered from object-centric neural fields, and motion generation is grounded by object-conditioned primitives combined over time.",{"name":82,"@type":73,"acceptedAnswer":83},"How does the temporal mixture-of-experts model produce full motion trajectories?",{"text":84,"@type":76},"A gating mechanism selects and blends multiple object-conditioned movement experts over time, composing primitives into complete trajectories while maintaining data efficiency.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,114,119,122,126,129,133],{"id":21,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":20,"doc_module":4,"doc_module_name":45,"category_name":111,"show_sort_weight":112,"slug":113},"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":28,"slug":125},9,"Religion & Spirituality","religion-spirituality",{"id":28,"doc_module":4,"doc_module_name":45,"category_name":127,"show_sort_weight":28,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":45,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":45,"category_name":135,"show_sort_weight":106,"slug":136},19,"General","general"]