[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-81569-en":3,"doc-seo-81569-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},81569,549758146520,"Patrick","https://ap-avatar.wpscdn.com/avatar/80002397d8c0411e94?_k=1775819394049821470",8,"Research & Report","Improving Language Agents through BREW: Bootstrapping Experientially-Learned Environmental Knowledge","Large Language Model (LLM)-based agents can perform complex multi-step tasks such as GUI automation, tool use, and data manipulation, yet they cannot learn from experience across sessions. BREW (Bootstrapping Experientially-Learned Environmental Knowledge) builds a structured, retrievable knowledge base of natural-language recipes that specify what to do, when to apply it, and what pitfalls to avoid. BREW decomposes memory into concept-localized documents and constructs the KB via state-space search using Expand-and-Gather MCTS, with hindsight relabeling to extract reusable competencies. Across OSWORLD, τ2-Bench, and SpreadSheetBench, BREW improves task success by 10–20% and reduces execution steps by 10–15%, while remaining inspectable, modular, and extensible.","arXiv :2511 .20297v2 [ cs .AI] 10 Jul 2026  \nImproving Language Agents through BREW: Bootstrapping expeRientially-learned Environmental knoWledge  \nShashank Kirtania1 ∗ Param Biyani2 Priyanshu Gupta3 Yasharth Bajpai3 Roshni Iyer4 Sumit Gulwani3 Gustavo Soares3  \n1University of Michigan, 2Independent, 3Microsoft, 4Apple.  \nAbstract  \nLarge Language Model (LLM)-based agents are increasingly capable of complex, multi-step tasks such as GUI automation, tool use, and data manipulation, yet they cannot learn from experience: each new session rediscovers solutions from scratch. We introduce BREW (Bootstrapping expeRientiallylearned Environmental knoWledge), a framework that distills an agent’s past interaction trajectories into a structured, retrievable knowledge base (KB) of natural-language recipes, concept-level procedural documents that capture what to do, when it applies, and what to watch out for. Drawing on the principle of library learning from program synthesis, BREW decomposes agent memory into modular, concept-localized documents and formalizes KB construction as a state-space search problem. To navigate this space, we introduce Expand-and-Gather Monte Carlo Tree Search (EGMCTS), a reward-guided algorithm that jointly optimizes recipe accuracy and retrievability across parallel, per-concept search trees. We further adapt hindsight relabeling to convert near-miss trajectories into positive demonstrations, surfacing latent agent competencies as reusable knowledge. On three domain-grounded benchmarks, OSWORLD, τ2-Bench, and SpreadSheetBench, BREW achieves 10–20% gains in task success and 10–15% fewer execution steps over base agents, while consistently outperforming existing memory-augmented baselines that can degrade below memoryless performance. The resulting KB is inspectable, modular, and extensible, providing a transparent and controllable substrate for agent optimization.  \n1 Introduction  \nLarge Language Model (LLM) based agents are increasingly capable of interacting with complex environments like navigating GUIs, calling external tools, and manipulating structured data across a wide range of real-world tasks (Li, 2025; Qin et al., 2025; Jimenez et al., 2024; Yang et al., 2024; Anthropic, 2024; OpenAI, 2025) . Yet these agents lack a fundamental ability that humans take for granted: learning from experience. When an agent encounters a familiar task in a new session, it rediscovers the solution from scratch, repeating the same wrong menu paths, redundant API calls, or failed spreadsheet manipulations it had already resolved before.  \nAfter exporting a document to PDF once in LibreOffice, a person does not memorize the raw click sequence; they internalize a recipe: the general pattern (File → Export as PDF), the conditions under which it applies (any Writer or Impress document), and the pitfalls to avoid (confirming the output opens correctly) . Similarly, after handling a few customer return requests, a support agent learns the procedural structure: authenticate the user, verify the order status is delivered, collect all exchange items upfront because the step cannot be repeated, then execute. These recipes transfer across tasks within the same environment, compressing future problem-solving into a few reliable steps.  \n∗1,2,4 Work done at Microsoft.  \nCorresponding author: [priyansgupta@microsoft.com](priyansgupta@microsoft.com)  \nspecific Grader  \nFigure 1: BREW architecture overview using examples from the OSWORLD dataset. Step 1 indicates trajectory generation with agent alignment to human-validated rubrics and correctness using a task-specific grader. Steps 2–4 indicate the Reflector Agent, which learns key concepts and insights from trajectories. Step 5 indicates the Integrator Agent, which integrates knowledge from the Reflector Agent to bootstrap the KB. We introduce Expand-and-Gather MCTS to find the best KB configuration by reward-guided search.  \nHow should an agent acquire such recipes? Weight optimization ","cbCailCGsQLPIoeK","https://ap.wps.com/l/cbCailCGsQLPIoeK","pdf",785062,4,1,41,"English","en",105,"# Abstract\n# Introduction\n## Learning recipes from experience\n## Limits of parameter baking and coarse memory\n## BREW and concept-level natural-language recipes\n## Expand-and-Gather MCTS for knowledge-base construction","[{\"question\":\"What problem does BREW address in LLM-based language agents?\",\"answer\":\"BREW targets the inability of agents to learn from experience across sessions, where each new session rediscovers solutions from scratch and repeats prior mistakes.\"},{\"question\":\"How does BREW represent experience as reusable knowledge?\",\"answer\":\"BREW distills past interaction trajectories into a knowledge base of natural-language recipes, where each recipe captures a concept-level procedural pattern, including applicability conditions and pitfalls.\"},{\"question\":\"How is the knowledge base constructed and optimized in BREW?\",\"answer\":\"BREW frames knowledge-base construction as a state-space search problem and uses Expand-and-Gather Monte Carlo Tree Search to jointly optimize recipe accuracy and retrievability across parallel per-concept search trees, with hindsight relabeling to convert near-miss trajectories into demonstrations.\"}]",1784174379,103,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"improving-language-agents-through-brew-bootstrapping-experientially-learned-environmental-knowledge","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/improving-language-agents-through-brew-bootstrapping-experientially-learned-environmental-knowledge/81569/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does BREW address in LLM-based language agents?","Question",{"text":75,"@type":76},"BREW targets the inability of agents to learn from experience across sessions, where each new session rediscovers solutions from scratch and repeats prior mistakes.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does BREW represent experience as reusable knowledge?",{"text":80,"@type":76},"BREW distills past interaction trajectories into a knowledge base of natural-language recipes, where each recipe captures a concept-level procedural pattern, including applicability conditions and pitfalls.",{"name":82,"@type":73,"acceptedAnswer":83},"How is the knowledge base constructed and optimized in BREW?",{"text":84,"@type":76},"BREW frames knowledge-base construction as a state-space search problem and uses Expand-and-Gather Monte Carlo Tree Search to jointly optimize recipe accuracy and retrievability across parallel per-concept search trees, with hindsight relabeling to convert near-miss trajectories into demonstrations.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]