[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-81783-en":3,"doc-seo-81783-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},81783,549758252649,"Ivy","https://ap-avatar.wpscdn.com/avatar/8000253669c5317157?_k=1778319167496531819",8,"Research & Report","Coachable Agents for Interactive Gameplay","Reinforcement learning supports advanced AI and robotics, but often converges to a single near-optimal behavior learned via trial and error. The work introduces a framework for coaching agents with controllable “styles” that modify how a core task is solved, ideally in real time. By combining universal value function approximators, curated training scenarios, and data augmentation, agents learn style-coherent behavior across complex domains while still completing the main task, enabling end-user selection at run time.","arXiv :2607 .00642v 1 [ cs .AI] 1 Jul 2026  \nCoachable agents for interactive gameplay  \nRoberto Capobianco 1 , Harm van Seijen2 , Nolan D. Bard2 , Neil Burch2 , Fatima Davelouis2 , Josh Davidson2 , Alisa Devlic 1 , Yunshu Du2 , Ishan Durugkar2 , Siddhant Gangapurwala2 , Daniel Hernandez2 , G. Zacharias Holland2 , Sahil Jain2 , Kenta Kawamoto3 , Raksha Kumaraswamy2 , Patrick MacAlpine2 , Dustin R. Morrill2 , Declan Oller2 , Francesco Riccio 1 , Akanksha Saran2 , Craig Sherstan3 , Kaushik Subramanian 1 , Thomas J. Walsh2 , Samuel Barrett2 , Kizza N. Frisbee2 , Mady Govil2 , Johannes Günther2 , Varun R. Kompella2 , James A. MacGlashan2 , Maxwell Svetlik2 , Michael D. Thomure2 , Jaden B. Travnik2 , Kevin Waugh2 , Elahe Aghapour2 , Florian Fuchs 1 , Andreanne Lemay2 , Shruti Mishra 1 , Takuma Seno3 , Peter Stone2 , Michael Spranger3 , and  \nPeter R. Wurman2,*  \n1 Sony AI, Zurich, Switzerland  \n2 Sony AI, North America (various locations)  \n3 Sony AI, Tokyo, Japan  \n* [peter.wurman@sony.com](peter.wurman@sony.com)  \nAbstract  \nReinforcement learning has proven to be a valuable tool in the creation of advanced AI and robotic systems, contributing to everything from game playing [46, 47, 61] to robotics [2, 30, 34] to foundation models [16, 41, 42, 49] . Through trial-and-error, these AI systems typically learn one, near-optimal behavior to solve their tasks. However, there are many use cases in which one would like to assert some level of control, preferably in real time, over how the task is solved. We refer to these modifications of a core task as styles. We combine universal value function approximators (UVFAs) with carefully selected training scenarios, learning algorithms, and data augmentation to create a framework for coaching agents that exhibit styles in complex domains. We demonstrate the framework’s application in the AAA video games Horizon Forbidden West and Gran Turismo, and in an open-source humanoid test domain. Despite the different nature of the domains—car racing, stylized game combat, and humanoid walking—each agent shows strong coherence to the style requests while still satisfying the main task in its domain. Importantly, the techniques outlined in this paper allow an end user to choose the final behavior at run time, giving them flexible control over the final executed performance.  \nIntroduction  \nWith artificial intelligence (AI) entering the mainstream, there is a growing focus on enabling users to meaningfully align the system’s behavior to the user’s preferences. While users are now familiar with requesting generative AI tools [3, 43] to follow style requests (\"generate an impressionist painting of a cat playing table tennis.\"), such control has so far been unavailable for agents operating in real-time control settings. Concurrently, reinforcement learning (RL) has become a leading method for achieving state-of-the-art performance across a wide range of real-time control domains [22, 28, 61] . These RL agents typically master a single task and perform it in a single, near-optimal way. In practice, in addition to what an agent accomplishes, users often care about how it accomplishes the task, which may depend  \non the context at run-time. For example, we might prefer our cleaning robot to clean up the room quietly because the baby is sleeping, or to prioritize speed because guests are about to arrive. These preferences do not change the underlying task itself—cleaning the room—but they do express valuable, context-dependent preferences over the agent’s behavior.  \nCoachable agents are trained to perform a task in a variety of different styles, and to apply those styleson-demand at execution time. This problem framing differs from multi-task[25, 50], multi-objective[1, 19] or goal-conditioned learning [32] in which the problem statement is focused on teaching an agent to accomplish multiple different tasks. To understand the difference, consider the robot cleaning scenario: classifying “cleaning quietly” an","cbCaihvMve2gpsfk","https://ap.wps.com/l/cbCaihvMve2gpsfk","pdf",9535634,2,1,38,"English","en",105,"# Abstract\n# Introduction\n# Approach","[{\"question\":\"What does the paper mean by “styles” for an agent?\",\"answer\":\"Styles are modifications of a core task that control secondary performance characteristics while keeping the main task the same. This lets users request context-dependent behavior, such as prioritizing quietness versus speed.\"},{\"question\":\"How does the proposed framework enable real-time user control of agent behavior?\",\"answer\":\"The method learns style-conditioned behavior so an end user can choose the final behavior at run time. The approach uses UVFAs plus carefully selected training scenarios and data augmentation to maintain coherence with style requests.\"},{\"question\":\"Which domains are used to demonstrate the framework’s effectiveness?\",\"answer\":\"The framework is demonstrated in AAA video games Horizon Forbidden West and Gran Turismo, and also in an open-source humanoid test domain. Despite domain differences, agents remain coherent to style requests while satisfying the primary task.\"}]",1784176117,96,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"coachable-agents-for-interactive-gameplay","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/coachable-agents-for-interactive-gameplay/81783/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What does the paper mean by “styles” for an agent?","Question",{"text":75,"@type":76},"Styles are modifications of a core task that control secondary performance characteristics while keeping the main task the same. This lets users request context-dependent behavior, such as prioritizing quietness versus speed.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the proposed framework enable real-time user control of agent behavior?",{"text":80,"@type":76},"The method learns style-conditioned behavior so an end user can choose the final behavior at run time. The approach uses UVFAs plus carefully selected training scenarios and data augmentation to maintain coherence with style requests.",{"name":82,"@type":73,"acceptedAnswer":83},"Which domains are used to demonstrate the framework’s effectiveness?",{"text":84,"@type":76},"The framework is demonstrated in AAA video games Horizon Forbidden West and Gran Turismo, and also in an open-source humanoid test domain. Despite domain differences, agents remain coherent to style requests while satisfying the primary task.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]