[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-81710-en":3,"doc-seo-81710-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},81710,3848291630094,"Emma Wilson","https://eur-avatar.wpscdn.com/davatar_085a072bc5b1113ac321206ff7593b45",8,"Research & Report","Libra: Training the Environment for Agentic Information Retrieval","Information localization within massive repositories is a cornerstone of agentic LLM systems. While synthetic-data optimization has advanced model training, optimizing the agent’s working environment—the repository itself—remains underexplored. Libra introduces self-evolving, mutable “catalogs” as hierarchical Markdown indices inside the repository. An LLM-driven loop uses a Prompter to generate synthetic queries, a frozen Solver to locate answers via catalogs, and a Healer to repair catalogs after failures, producing continual logarithmic gains in localization accuracy.","Libra: Training the Environment for Agentic Information Retrieval  \nXuan Zhao  \n[xuan.zhao@salesforce.com](xuan.zhao@salesforce.com)  \nAndy Chiu  \n[andy.chiu@salesforce.com](andy.chiu@salesforce.com)  \narXiv :2607 .00016v1 [ cs .IR] 26 May 2026  \nGengyu Wang∗  \n[gengyu.wang@columbia.edu](gengyu.wang@columbia.edu)  \nAbstract  \nInformation localization within massive repositories is a cornerstone of agentic LLM systems. While synthetic data-driven optimization has proven successful in training LLMs, little attention has been paid to optimizing the agent’s working environment (the repository itself) in a data-driven manner. To bridge this gap, we present LIBRA, a self-evolving framework that introduces mutable “catalogs”  \n(hierarchical Markdown files serving as navigable indices) into the repository. LIBRA runs an LLM-driven optimization loop where a Prompter generates synthetic queries, a frozen Solver attempts to resolve them by navigating the catalogs, and a Healer rewrites the catalogs in response to the Solver’s localization failures. Evaluations across 12 SWE-BENCH LITE repositories demonstrate that this environmental healing yields continual, logarithmic improvements in code localization accuracy.  \nFurthermore, these environmental improvements transfer zero-shot across different LLMs and problem sets. Although the focus of this paper is to study the general behavior of such a system, we also demonstrate that a minimalist coding agent equipped with LIBRA-optimized catalogs outperforms state-of-the-art baselines.  \nCode is available at [https://github.com/salesforce-misc/Libra](https://github.com/salesforce-misc/Libra and data)[ and data](https://github.com/salesforce-misc/Libra and data)  \nat [https://huggingface.co/datasets/Salesforce/Libra](https://huggingface.co/datasets/Salesforce/Libra).  \n1 Introduction  \nCode localization within large repositories is a fundamental capability for autonomous coding agents [Anthropic, 2024, Anysphere, 2024, Yang et al., 2024b, Wang et al., 2024] . Localization accuracy has further been shown to correlate strongly with downstream task resolution rates [Chenet al., 2025, Wang et al., 2026], making it one of the key contributors to end-to-end agent performance. More broadly, information retrieval serves as the cornerstone of agentic systems and has consequently been the subject of extensive research [Lewis et al., 2020, Jin et al., 2025, Zhang et al., 2026] . These agentic systems generally operate on three interacting pillars: the underlying foundation model, the agent design (prompt engineering and tool use), and the environment itself (repository contents, search indices, and documentation) .  \nTo improve an agent’s ability to navigate a specific environment, a natural approach is to fine-tune the underlying foundation model using repository-derived data [Jimenez et al., 2024, Ma et al., 2024, Pan et al., 2025, Wei et al., 2025] . While this allows the model to internalize environmental knowledge into its weights using mature optimization techniques, it inherently locks the system to a specific model. Consequently, the agent may lose access to the superior reasoning and semantic capabilities of state-of-the-art commercial LLMs, rendering this approach sub-optimal for practical, evolving applications.  \n∗ Corresponding author.  \nPreprint.  \nFigure 1: Overview of the Libra system. Three frozen agents (Prompter, Solver, Healer) drive an adversarial loop in which only the Markdown catalog is updated. The Prompter sees a random file chunk and fabricates a question whose answer is that chunk; the Solver does not see the chunk and must answer by searching the repository alongside the catalog; the Healer reads the accumulated failure report once per batch and edits the catalog to repair the routing signal.  \nAlternatively, researchers have sought to augment the agent’s environment by constructing static or reverse indices [Jin et al., 2025, Chen et al., 2025, Wang et al., 2026, Zhang et al., ","cbCaio1Ip8RHOfKt","https://ap.wps.com/l/cbCaio1Ip8RHOfKt","pdf",1530381,3,1,21,"English","en",105,"# Abstract\n# Introduction\n## Agentic LLM systems and information retrieval\n## Limitations of model fine-tuning and static indices\n## Interaction-trace and memory-based approaches\n## Libra framework and adversarial optimization loop\n## Evaluation on SWE-BENCH LITE","[{\"question\":\"What problem does LIBRA address in agentic LLM systems?\",\"answer\":\"LIBRA targets improving information localization in large repositories, focusing on training the agent’s environment (the repository index) rather than only optimizing the foundation model.\"},{\"question\":\"How does the LIBRA training loop work?\",\"answer\":\"A Prompter generates synthetic queries from random code chunks, a frozen Solver answers by navigating the Markdown catalogs, and a Healer edits the catalogs based on the Solver’s localization failures.\"},{\"question\":\"What evidence does the document provide for LIBRA’s effectiveness?\",\"answer\":\"Evaluations across 12 SWE-BENCH LITE repositories show continual, logarithmic improvements in code localization accuracy, and the gains transfer zero-shot to different LLMs and problem sets.\"}]",1784175564,53,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"libra-training-the-environment-for-agentic-information-retrieval","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/libra-training-the-environment-for-agentic-information-retrieval/81710/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does LIBRA address in agentic LLM systems?","Question",{"text":75,"@type":76},"LIBRA targets improving information localization in large repositories, focusing on training the agent’s environment (the repository index) rather than only optimizing the foundation model.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the LIBRA training loop work?",{"text":80,"@type":76},"A Prompter generates synthetic queries from random code chunks, a frozen Solver answers by navigating the Markdown catalogs, and a Healer edits the catalogs based on the Solver’s localization failures.",{"name":82,"@type":73,"acceptedAnswer":83},"What evidence does the document provide for LIBRA’s effectiveness?",{"text":84,"@type":76},"Evaluations across 12 SWE-BENCH LITE repositories show continual, logarithmic improvements in code localization accuracy, and the gains transfer zero-shot to different LLMs and problem sets.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]