[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-seo-141363-105":3,"detail-sidebar-cat-0-en-105":81,"doc-detail-141363-en":130},{"code":4,"msg":5,"data":6},0,"ok",{"site_id":7,"language":8,"slug":9,"title":10,"keywords":11,"description":12,"schema_data":13,"social_meta":74,"head_meta":76,"extra_data":78,"updated_unix":80},105,"en","reinforcement-learning-with-model-based-feedforward-inputs-for-robotic-table-tennis","Reinforcement learning with model-based feedforward inputs for robotic table tennis","","The work reframes reinforcement learning by optimizing feedforward inputs rather than feedback policies, aiming to stabilize training and shift most learning effort toward a supervised learning formulation. Labels are generated using a variant of iterative learning control that incorporates prior knowledge about the robot’s dynamics. The framework is evaluated on real-world ping-pong interception and returning with a four-degree-of-freedom pneumatic arm, where control and learning are challenging. Compared with feedback-policy optimization, it improves success rate (100% vs. 96% over 107 trials) and reduces required training samples by about a factor of ten while handling varying incoming trajectories.",{"@graph":14,"@context":73},[15,34,56],{"@type":16,"itemListElement":17},"BreadcrumbList",[18,23,27,31],{"item":19,"name":20,"@type":21,"position":22},"https://docshare.wps.com","Home","ListItem",1,{"item":24,"name":25,"@type":21,"position":26},"https://docshare.wps.com/document/","Document",2,{"item":28,"name":29,"@type":21,"position":30},"https://docshare.wps.com/document/research-report/","Research & Report",3,{"item":32,"name":10,"@type":21,"position":33},"https://docshare.wps.com/document/reinforcement-learning-with-model-based-feedforward-inputs-for-robotic-table-tennis/141363/",4,{"url":32,"name":10,"@type":35,"image":36,"author":41,"headline":10,"publisher":44,"fileFormat":47,"inLanguage":8,"description":12,"dateModified":48,"datePublished":49,"encodingFormat":47,"isAccessibleForFree":50,"interactionStatistic":51},"DigitalDocument",{"url":37,"@type":38,"width":39,"height":40},"https://docshare.wps.com/thumbnails/reinforcement-learning-with-model-based-feedforward-inputs-for-robotic-table-tennis/141363.png","ImageObject",300,407,{"name":42,"@type":43},"วิน","Person",{"url":19,"name":45,"@type":46},"DocShare","Organization","application/pdf","2026-10-06","2026-08-25",true,{"@type":52,"interactionType":53,"userInteractionCount":55},"InteractionCounter",{"@type":54},"ViewAction",6,{"@type":57,"mainEntity":58},"FAQPage",[59,65,69],{"name":60,"@type":61,"acceptedAnswer":62},"What key change does the framework make compared with traditional reinforcement learning?","Question",{"text":63,"@type":64},"It optimizes over feedforward inputs instead of optimizing over feedback policies, which reduces the risk of destabilizing the system during training.","Answer",{"name":66,"@type":61,"acceptedAnswer":67},"How are the training labels generated in this approach?",{"text":68,"@type":64},"Labels are generated with a variant of iterative learning control that also incorporates prior knowledge about the underlying robot dynamics.",{"name":70,"@type":61,"acceptedAnswer":71},"How does the method perform in real-world robotic table tennis experiments?",{"text":72,"@type":64},"The framework achieves a higher return success rate than a feedback-policy reinforcement learning approach (100% vs. 96% across 107 consecutive trials) and needs about one tenth of the training samples.","https://schema.org",{"og:url":32,"og:type":75,"og:title":10,"og:site_name":45,"og:description":12},"article",{"robots":77,"canonical":32},"index,follow",{"doc_id":79,"site_id":7},141363,1787654928,{"code":4,"msg":82,"data":83},"success",[84,88,92,96,101,105,110,114,119,122,126],{"id":22,"doc_module":4,"doc_module_name":25,"category_name":85,"show_sort_weight":86,"slug":87},"Story & Novel",90,"story-novel",{"id":26,"doc_module":4,"doc_module_name":25,"category_name":89,"show_sort_weight":90,"slug":91},"Literature",80,"literature",{"id":33,"doc_module":4,"doc_module_name":25,"category_name":93,"show_sort_weight":94,"slug":95},"Exam",70,"exam",{"id":97,"doc_module":4,"doc_module_name":25,"category_name":98,"show_sort_weight":99,"slug":100},5,"Comic",60,"comic",{"id":55,"doc_module":4,"doc_module_name":25,"category_name":102,"show_sort_weight":103,"slug":104},"Technology",50,"technology",{"id":106,"doc_module":4,"doc_module_name":25,"category_name":107,"show_sort_weight":108,"slug":109},7,"Healthcare",40,"healthcare",{"id":111,"doc_module":4,"doc_module_name":25,"category_name":29,"show_sort_weight":112,"slug":113},8,30,"research-report",{"id":115,"doc_module":4,"doc_module_name":25,"category_name":116,"show_sort_weight":117,"slug":118},9,"Religion & Spirituality",20,"religion-spirituality",{"id":117,"doc_module":4,"doc_module_name":25,"category_name":120,"show_sort_weight":117,"slug":121},"World Cup","world-cup",{"id":123,"doc_module":4,"doc_module_name":25,"category_name":124,"show_sort_weight":123,"slug":125},10,"Lifestyle","lifestyle",{"id":127,"doc_module":4,"doc_module_name":25,"category_name":128,"show_sort_weight":97,"slug":129},19,"General","general",{"code":4,"msg":82,"data":131},{"doc_id":79,"user_id":132,"nickname":42,"user_avatar":133,"doc_module":4,"category_id":111,"category_name":29,"doc_title":10,"doc_description":12,"doc_content":134,"file_id":135,"file_url":136,"file_type":137,"file_size":138,"view_count":55,"is_deleted":4,"is_public":22,"is_downloadable":22,"audit_status":22,"page_count":139,"language":140,"language_code":8,"site_id":7,"html_lang":8,"table_of_contents":141,"faqs":142,"seo_title":143,"seo_description":12,"update_tm":80,"read_time":144},2336475104736,"https://ap-avatar.wpscdn.com/avatar/22000c4c5e0e5b17e70?x-image-process=image/resize,m_fixed,w_180,h_180&k=1786591360781797222","Reinforcement learning with model-based feedforward inputs for robotic table tennis  \nHao Ma1 · Dieter Büchler2 · Bernhard Schölkopf2 · Michael Muehlebach1  \nReceived: 31 January 2023 / Accepted: 18 September 2023 / Published online: 17 October 2023 © The Author(s) 2023  \nAbstract  \nWe rethink the traditional reinforcement learning approach, which is based on optimizing over feedback policies, and propose a new framework that optimizes over feedforward inputs instead. This not only mitigates the risk of destabilizing the system during training but also reduces the bulk of the learning to a supervised learning task. As a result, efﬁcient and well-understood supervised learning techniques can be applied and are tuned using a validation data set. The labels are generated with a variant ofiterative learning control, which also includes prior knowledge about the underlying dynamics. Our framework is applied for intercepting and returning ping-pong balls that are played to a four-degrees-of-freedom robotic arm in real-world experiments. The robot arm is driven by pneumatic artiﬁcial muscles, which makes the control and learning tasks challenging. We highlight the potential of our framework by comparing it to a reinforcement learning approach that optimizes over feedback policies. We ﬁnd that our framework achieves a higher success rate for the returns (100% vs. 96%, on 107 consecutive trials, see [https://youtu.be/kR9jowEH7PY](https://youtu.be/kR9jowEH7PY)) while requiring only about one tenth of the samples during training. We also ﬁnd that our approach is able to deal with a variant of different incoming trajectories.  \nKeywords Reinforcement learning · Iterative learning control · Supervised learning · Table tennis robot · Pneumatic artiﬁcial muscle · Soft robotics  \n1 Introduction  \nReinforcement learning has been proven to be highly effective in a variety of contexts. An important example is AlphaGo Zero (Silver et al., 2016, 2017), which managed to completely overpower all human players in the game of Go. Other examples include the work of Oh et al.(2016), Tessleretal.(2017), Firoiuetal.(2017), Kanskyetal.(2017) that focused on video games, the work of Yogatama  \nB Hao Ma[hao.ma@tuebingen.mpg.de](hao.ma@tuebingen.mpg.de)  \nDieter Büchler  \n[dieter.buechler@tuebingen.mpg.de](dieter.buechler@tuebingen.mpg.de)  \nBernhard Schölkopf  \n[bernhard.schoelkopf@tuebingen.mpg.de](bernhard.schoelkopf@tuebingen.mpg.de)  \nMichael Muehlebach  \n[michael.muehlebach@tuebingen.mpg.de](michael.muehlebach@tuebingen.mpg.de)  \n1 Learning and Dynamical Systems, Max Planck Institute for Intelligent Systems, Max-Planck-Ring 4, Tübingen 72076, Germany  \n2 Empirical Inference, Max Planck Institute for Intelligent Systems, Max-Planck-Ring 4, Tübingen 72076, Germany  \net al. (2016), Paulus et al. (2017), Zhang and Lapata (2017) that focused on natural language processing, and the work of Liu et al. (2017), Devrim Kaba et al. (2017), Cao et al. (2017), Brunner et al. (2018) that focused on computer vision. Despite these successes, where reinforcement learning agents are shown to compete and outperform humans, researchers have struggled to achieve a similar level of success in robotics applications. We identify the following key bottlenecks, which we believe hinder the application of reinforcement learning to robotic systems. This also motivatesour work, which proposes a new reinforcement learning scheme that addresses some of these shortcomings.  \nThe ﬁrst factor is that the lack of prior knowledge causes reinforcement learning algorithms to sometimes apply relatively aggressive feedback policies 1 during training. This has the potential to cause irreversible damage to robotic systems, which are often expensive and require careful maintenance (Moldovan & Abbeel, 2012; Schneider, 1996) . Moreover, these aggressive policies are typically not effec-  \n1 The term“aggressive feedback policy\"refers toa policy that generates actions with the potential to induce har","cbCaimeTWEIIBVF2","https://ap.wps.com/l/cbCaimeTWEIIBVF2","pdf",1446455,17,"English","# Abstract\n# Introduction\n## Bottlenecks in robotic reinforcement learning\n## Key motivation for the proposed feedforward framework","[{\"question\":\"What key change does the framework make compared with traditional reinforcement learning?\",\"answer\":\"It optimizes over feedforward inputs instead of optimizing over feedback policies, which reduces the risk of destabilizing the system during training.\"},{\"question\":\"How are the training labels generated in this approach?\",\"answer\":\"Labels are generated with a variant of iterative learning control that also incorporates prior knowledge about the underlying robot dynamics.\"},{\"question\":\"How does the method perform in real-world robotic table tennis experiments?\",\"answer\":\"The framework achieves a higher return success rate than a feedback-policy reinforcement learning approach (100% vs. 96% across 107 consecutive trials) and needs about one tenth of the training samples.\"}]","Reinforcement learning with model-based feedforward inputs for robotic table tennis | PDF",43]