[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85348-en":3,"doc-seo-85348-105":29,"detail-sidebar-cat-0-en-105":90},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":11,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},85348,13056703020460,"Valentina","https://ap-avatar.wpscdn.com/avatar/be000253dac470eee5d?_k=1778207105932848923",8,"Research & Report","SKooP Symmetric Koopman Predictions for Faster and More Generalizable Legged Robot Locomotion with Reinforcement Learning","Reinforcement learning for robotic control often suffers from poor sample efficiency, while physics-informed methods are frequently demonstrated only on low-dimensional benchmarks. SKooP (Symmetric Koopman Predictions) targets high-dimensional legged robots by learning a Koopman model of system dynamics alongside the policy using an autoencoder. Koopman predictions provide privileged critic observations for smoother features, and group symmetries are integrated into actor, critic, encoder, and decoder to yield a highly equivariant policy. Experiments show reduced convergence time, higher rewards, and policy transfer across simulation environments on challenging bipedal locomotion tasks.","SKooP: Symmetric Koopman Predictions for Faster and More Generalizable Legged Robot Locomotion with Reinforcement Learning  \nEvelyn D’Elia 1 ,2 , Weishu Zhan2 , Giulio Turrisi3 , Giulio Romualdi4 , Giuseppe L’Erario4 , Raffaello Camoriano5 ,6 , Wei Pan7 , and Daniele Pucci4  \narXiv :2607 . 11624v1 [ cs .RO] 13 Jul 2026  \nAbstract—Reinforcement learning (RL) algorithms classically suffer from poor sample efficiency. In robotics, a recent line of work has emerged addressing this problem by encoding physics priors in the learning process. However, most of these approaches are validated on well-defined, low-dimensional benchmark systems rather than high-dimensional robots with complex nonlinear dynamics. In this paper, we introduce SKooP (Symmetric Koopman Predictions), an approach combining the advantages of morphological symmetries with those of a Koopman model learned via autoencoder to enhance policy learning. SKooP learns a Koopman model of the system dynamics alongside the policy. The resulting Koopman predictions are used as privileged observations for the critic, allowing the agent to learn based on smoother, more informative features. We also incorporate group symmetries into the actor, critic, encoder and decoder networks to produce a highly equivariant policy. The SKooP approach is validated via in-depth analysis of the learned Koopman models and symmetric policies to showcase how each of these influences the agent’s performance. We also show that the learned policies are transferable to different simulation environments. Our results show that SKooP consistently reduces convergence time and increases the learned reward for multiple challenging bipedal locomotion tasks on a quadruped robot. Project page: [https://evelyd.github](https://evelyd.github) . io/SymmetricKoopmanPredictions/  \nI. INTRODUCTION  \nThe use of reinforcement learning approaches for robotic control has skyrocketed in recent years, mainly thanks to advances in computational efficiency and simulation fidelity. However, for systems such as legged robots which have complex, nonlinear dynamics, learning effective policies is still a challenging and active research area. Although powerful, purely data-driven model-free RL approaches rely on costly  \n*This study was carried out within the FAIR - Future Artificial Intelligence Research and received funding from the European Union Next-GenerationEU (PIANO NAZIONALE DI RIPRESA E RESILIENZA (PNRR) – MISSIONE 4 COMPONENTE 2, INVESTIMENTO 1 .3 – D.D. 1555 11/10/2022, PE00000013) . This manuscript reflects only the authors’ views and opinions, neither the European Union nor the European Commission can be considered responsible for them.  \n1IIT@MIT, Italian Institute of Technology (IIT), 16163 Genoa, Italy evelyn .delia@iit .it  \n2Machine Learning and Optimisation, University of Manchester, M13 9PL Manchester, U.K.  \n3Dynamic Legged Systems Laboratory, IIT, 16163 Genoa, Italy  \n4 Generative Bionics S.R.L, 16163 Genoa, Italy  \n5DAUIN, Politecnico di Torino, 10129 Turin, Italy  \n6Rehab Technologies Lab, IIT, 16163 Genoa, Italy  \n7 School of Engineering, Newcastle University, NE1 7RU Newcastle upon Tyne, U.K.  \nE. D’Elia, G. Romualdi, G. L’Erario, and D. Pucci contributed to this work while at the Artificial and Mechanical Intelligence lab, IIT, Italy.  \nW. Pan contributed to this work while at the Machine Learning and Optimisation group, University of Manchester, U.K.  \nFig. 1. Comparison of SKooP performance on trained vs. mirrored push door task. Top row: right-opening door (training task) . Bottom row: leftopening door (unseen symmetric task) .  \ntrial and error. Conversely, more traditional model-based control approaches exploit the known physics of the system, but typically require expert knowledge and extensive manual tuning. Fusing model-based control with RL holds promise for overcoming their respective limitations and devising more efficient and effective robot control methods.  \nModel-based control methods em","cbCaicdC86C6WHny","https://ap.wps.com/l/cbCaicdC86C6WHny","pdf",1183451,2,1,"English","en",105,"# Introduction\n## Sample efficiency challenges in model-free RL\n## Physics priors via symmetry and Koopman linearization","[{\"question\":\"What problem does SKooP address in reinforcement learning for legged robots?\",\"answer\":\"SKooP targets poor sample efficiency and limited generalization typical of model-free RL, especially on robots with complex nonlinear, high-dimensional dynamics.\"},{\"question\":\"How does SKooP use Koopman predictions during learning?\",\"answer\":\"SKooP learns a Koopman model of system dynamics alongside the policy, then uses Koopman predictions as privileged observations for the critic so the agent learns from smoother, more informative features.\"},{\"question\":\"What role do symmetries play in SKooP?\",\"answer\":\"SKooP incorporates group symmetries into the actor, critic, encoder, and decoder networks to produce a highly equivariant policy that improves learning efficiency and generalization.\"}]",1784202698,20,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":85,"head_meta":87,"extra_data":89,"updated_unix":27},"skoop-symmetric-koopman-predictions-for-faster-and-more-generalizable-legged-robot-locomotion-with-reinforcement-learning","",{"@graph":35,"@context":84},[36,52,67],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,46,49],{"item":40,"name":41,"@type":42,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":20},"https://docshare.wps.com/document/","Document",{"item":47,"name":12,"@type":42,"position":48},"https://docshare.wps.com/document/research-report/",3,{"item":50,"name":13,"@type":42,"position":51},"https://docshare.wps.com/document/skoop-symmetric-koopman-predictions-for-faster-and-more-generalizable-legged-robot-locomotion-with-reinforcement-learning/85348/",4,{"url":50,"name":13,"@type":53,"author":54,"headline":13,"publisher":56,"fileFormat":59,"inLanguage":23,"description":14,"dateModified":60,"datePublished":61,"encodingFormat":59,"isAccessibleForFree":62,"interactionStatistic":63},"DigitalDocument",{"name":9,"@type":55},"Person",{"url":40,"name":57,"@type":58},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":64,"interactionType":65,"userInteractionCount":20},"InteractionCounter",{"@type":66},"ViewAction",{"@type":68,"mainEntity":69},"FAQPage",[70,76,80],{"name":71,"@type":72,"acceptedAnswer":73},"What problem does SKooP address in reinforcement learning for legged robots?","Question",{"text":74,"@type":75},"SKooP targets poor sample efficiency and limited generalization typical of model-free RL, especially on robots with complex nonlinear, high-dimensional dynamics.","Answer",{"name":77,"@type":72,"acceptedAnswer":78},"How does SKooP use Koopman predictions during learning?",{"text":79,"@type":75},"SKooP learns a Koopman model of system dynamics alongside the policy, then uses Koopman predictions as privileged observations for the critic so the agent learns from smoother, more informative features.",{"name":81,"@type":72,"acceptedAnswer":82},"What role do symmetries play in SKooP?",{"text":83,"@type":75},"SKooP incorporates group symmetries into the actor, critic, encoder, and decoder networks to produce a highly equivariant policy that improves learning efficiency and generalization.","https://schema.org",{"og:url":50,"og:type":86,"og:title":13,"og:site_name":57,"og:description":14},"article",{"robots":88,"canonical":50},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":91},[92,96,100,104,109,114,119,122,126,129,133],{"id":21,"doc_module":4,"doc_module_name":45,"category_name":93,"show_sort_weight":94,"slug":95},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":45,"category_name":97,"show_sort_weight":98,"slug":99},"Literature",80,"literature",{"id":51,"doc_module":4,"doc_module_name":45,"category_name":101,"show_sort_weight":102,"slug":103},"Exam",70,"exam",{"id":105,"doc_module":4,"doc_module_name":45,"category_name":106,"show_sort_weight":107,"slug":108},5,"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":45,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":28,"slug":125},9,"Religion & Spirituality","religion-spirituality",{"id":28,"doc_module":4,"doc_module_name":45,"category_name":127,"show_sort_weight":28,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":45,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":45,"category_name":135,"show_sort_weight":105,"slug":136},19,"General","general"]