[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83341-en":3,"doc-seo-83341-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83341,687197207919,"Theodora","https://ap-avatar.wpscdn.com/avatar/a000253d6f5f7c60be?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779446848396160552",8,"Research & Report","Open-ended Multi-agent Autocurricula via Visual Inspection of Policies","Open-ended reinforcement learning curricula aim to train generally capable agents by selecting tasks that progressively unlock more complex skills. A core difficulty lies in judging task difficulty against the agent’s current learning progress. Prior methods use scalar task scores or text summaries, which can miss subtle progress when rewards are sparse or misleading. This work introduces Visual Inspection of Policies (VIP), using a multi-modal video language model to inspect policy videos and recommend next tasks, improving curricula on SMAC versus text-only and score-based baselines.","OPEN-ENDED MULTI-AGENT AUTOCURRICULA VIA VISUAL INSPECTION OF POLICIES WITH MULTI-MODAL LLMS  \nLorenzo Pant1 Andrea Fanti1 Roberto Capobianco2  \n1 Sapienza University of Rome, Italy 2 Sony AI, Zurich, Switzerland [pante.1885460@studenti.uniroma1.it](pante.1885460@studenti.uniroma1.it) [fanti@diag.uniroma1.it](fanti@diag.uniroma1.it)[ ](fanti@diag.uniroma1.it)[roberto.capobianco@sony.com](roberto.capobianco@sony.com)  \narXiv :2607 .08 193v 1 [ cs .LG] 9 Jul 2026  \nABSTRACT  \nOpen-ended curricula in Reinforcement Learning (RL) aim to train generally-capable agents by identifying tasks that facilitate learning increasingly complex skills. A major challenge when designing such curricula is assessing task difficulty relative to the agent’s current learning progress. While previous work has explored using scalar task scores or textual summaries of the agent’s behavior, here we study a different approach: directly inspecting policy behavior via recorded episode videos. We introduce a simple yet effective instantiation of this approach which leverages a Video Language Model (VLM) to both process these videos and provide curriculum recommendations, which we call Visual Inspection of Policies (VIP) . Since videos can naturally contain any number of controllable agents, we empirically study VIP on the StarCraft Multi-Agent Challenge (SMAC) . We show that even with a lightweight and openly accessible VLM (VideoLLaMa2-7B), VIP can use policy videos to generate more effective curricula than both its text-only ablation and methods that rely on scalar task scores.  \nTeam Member Panics  \n(a) (b) (c) (d)  \nFigure 1: An example of how policy videos can capture promising curriculum directions that are inaccessible to methods based on learning signal. The four images are selected frames from an episode video recorded in the StarCraft Multi-Agent Challenge (SMAC) environment. While allied troops (green) have obtained a clear numbers advantage with an effective strategy in (a), they still end up losing this match to a single enemy (red) after panicking in (b) and (d) . Despite the 0% win rate, agents are close to discovering a winning strategy—notably, this insight is inaccessible to any curriculum curator that relies on local learning signals (e.g. Unsupervised Environment Design (Dennis et al., 2020)) . Here we instead propose a method to harness this information with a Vision Language Model (VLM), which we call Visual Inspection of Policies. In this exact scenario, VIP chooses to continue training on the same task, achieving a substantial improvement to an ∼80% win rate. Quantitative results supporting this intuition are reported in Section 5 (text-only ablation); other related qualitative analyses are also reported in Appendix D.  \n1 INTRODUCTION  \nMany real-world Reinforcement Learning (RL) tasks are difficult or impossible to solve by simply starting from a random policy and directly trying to optimize its downstream performance (Wang et al., 2019 ; 2020a) . These tasks typically require intermediate “stepping stone” skills that are hard to discover with direct optimization and are impossible to identify in advance without extensive domain knowledge. This has motivated research into methods to automatically design open-ended curricula (Wang et al., 2020b ; Dennis et al., 2020 ; Zhang et al., 2024a ; Wang et al., 2023) which both discover and solve increasingly complex tasks. The ultimate goal of these methods is to produce generally-capable agents that can quickly adapt to any task in a particular domain, including those unseen in training.  \nTextual Summary  \nPolicy Video  \nVideo Language Model  \nRaw Task Recommendation  \nTask Space  \nSentence Similarity  \nNext Interesting Task  \nFigure 2: An overview of Visual Inspection of Policies (VIP) . At each step of the curriculum, the agent is trained on the current task with a (Multi-Agent) Reinforcement Learning algorithm. Then, one or more recordings of episodes from the current policy are fed to ","cbCaibOfEP1HPked","https://ap.wps.com/l/cbCaibOfEP1HPked","pdf",3552434,4,1,22,"English","en",105,"# Introduction\n## Visual Inspection of Policies (VIP)\n## Curriculum difficulty tradeoff\n## Limitations of scalar scores and text-only summaries","[{\"question\":\"What problem does open-ended multi-agent autocurricula address?\",\"answer\":\"It targets reinforcement learning settings where starting from a random policy is insufficient and intermediate “stepping stone” skills must be discovered automatically through an evolving curriculum.\"},{\"question\":\"How does VIP recommend the next curriculum task?\",\"answer\":\"VIP records episodes from the current policy and feeds the policy videos, along with a small textual summary (e.g., win rate), into a video language model to recommend the next interesting task.\"},{\"question\":\"Why do scalar task scores or text-only summaries fall short?\",\"answer\":\"Scalar learning signals can fail when rewards are sparse or deceptive, and text-only summaries may not capture progress that is evident from visual behavior cues in the policy execution.\"}]",1784186876,55,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"open-ended-multi-agent-autocurricula-via-visual-inspection-of-policies","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/open-ended-multi-agent-autocurricula-via-visual-inspection-of-policies/83341/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does open-ended multi-agent autocurricula address?","Question",{"text":75,"@type":76},"It targets reinforcement learning settings where starting from a random policy is insufficient and intermediate “stepping stone” skills must be discovered automatically through an evolving curriculum.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does VIP recommend the next curriculum task?",{"text":80,"@type":76},"VIP records episodes from the current policy and feeds the policy videos, along with a small textual summary (e.g., win rate), into a video language model to recommend the next interesting task.",{"name":82,"@type":73,"acceptedAnswer":83},"Why do scalar task scores or text-only summaries fall short?",{"text":84,"@type":76},"Scalar learning signals can fail when rewards are sparse or deceptive, and text-only summaries may not capture progress that is evident from visual behavior cues in the policy execution.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]