[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83059-en":3,"doc-seo-83059-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83059,13056703019404,"Miles","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","Information Gain-based Rollout Policy Optimization","Reinforcement learning for large language model (LLM) agents on long-horizon search tasks faces a limitation: rollout budgets are allocated without explicitly measuring the utility of intermediate states. This can waste computation on low-value branches despite large differences in informativeness. IGRPO introduces information-gain-based, budget-aware tree-structured rollouts that expand nodes proportional to node-level informativeness, suppressing unpromising branches. Experiments on seven search-augmented QA benchmarks show consistent improvements under equal rollout budgets by leveraging an induced teacher distribution for principled policy optimization.","arXiv :2607 .06223v 1 [ cs .AI ] 7 Jul 2026  \nInformation Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents  \nYijun Zhang∗ Fan Xu∗ Jiaxin Ding† Yule Xie Shiqing Gao  \nXin Ding Haoxiang Zhang Luoyi Fu Xinbing Wang Shanghai Jiao Tong University  \n∗Equal contribution. †Corresponding author.  \nAbstract  \nReinforcement learning has become a promising paradigm for improving large language model (LLM) agents on long-horizon search tasks, where the agent must make a sequence of intermediate decisions before receiving a final outcome.  \nHowever, existing methods still face a key limitation: the rollout budget is often allocated without explicitly assessing the utility of intermediate states. As a result, substantial computation may be spent on low-value states, even though different branches can vary drastically in their informativeness. In this paper, we propose Information Gain-based Rollout Policy Optimization (IGRPO), a policy optimization framework that treats intermediate-state informativeness as the organizing principle of rollout collection. Specifically, IGRPO performs budget-aware tree-structured rollouts by allocating expansion budget according to node-level informativeness, so that more informative branches are expanded more frequently while unpromising branches are progressively suppressed. We further demonstrate that the information gain-based rollout induces an explicit limiting teacher distribution over trajectories, which naturally yields a clear policy optimization target, thereby unifying adaptive tree-structured exploration with principled policy learning under a single framework. Experiments on seven challenging search-augmented QA benchmarks demonstrate that IGRPO consistently outperforms strong baselines under the same rollout budget constraints, validating the effectiveness of leveraging the induced teacher distribution to guide policy optimization for long-horizon search agents.  \n1 Introduction  \nLarge language models [1–3] are increasingly trained to act as agents that solve tasks through multi-turn interaction with external tools [4–7] . In search-augmented question answering [8–10], an agent must decide what to reason about, when to issue a retrieval query, how to use the returned evidence, and finally when to answer. A central difficulty in this setting is that the agent must allocate a limited interaction budget across many possible intermediate search states. Effective training therefore requires not only assigning credit to completed trajectories, but also deciding where rollout computation should be spent.  \nReinforcement learning has been widely adopted to improve such agents by optimizing their interaction trajectories [11] . Existing outcome-based methods typically optimize complete trajectories using final correctness rewards [12, 13], while group-based variants such as GRPO [14] further avoid a learned critic by comparing multiple rollouts for the same question. Recent studies have sought to improve long-horizon agent training from two complementary directions: finer-grained credit assignment and broader exploration. Chain-based methods such as GiGPO [15] and IGPO [16] move beyond pure outcome-level learning by introducing turn-level learning signals for intermediate  \nPreprint. Code is available at [https://github.com/e3trange/IGRPO](https://github.com/e3trange/IGRPO).  \nFigure 1: Illustration of different rollout patterns. Red nodes represent unpromising or even misleading states, green nodes represent informative states, and gray nodes represent other intermediate states. Left: both chain-based and existing tree-based methods may still allocate exploration budget to unpromising nodes. Right: our method adopts information gain-based rollout allocation, preferentially expanding informative nodes and reducing exploration over unpromising branches.  \ndecisions. Tree-based methods such as Tree-GRPO [17] and AEPO [18] extend rollout generat","cbCaicmu6sY8rJLw","https://ap.wps.com/l/cbCaicmu6sY8rJLw","pdf",3958403,3,1,19,"English","en",105,"# Introduction\n## Long-horizon search agents and budget allocation\n## Related work: outcome-based, group-based, and chain/tree methods\n## Information gain-based rollout policy optimization (IGRPO)","[{\"question\":\"What problem does IGRPO address in long-horizon LLM search training?\",\"answer\":\"It targets inefficient rollout budget allocation caused by not explicitly estimating how useful intermediate states are, leading to wasted computation on low-value or misleading branches.\"},{\"question\":\"How does IGRPO decide which tree nodes to expand during rollouts?\",\"answer\":\"IGRPO treats intermediate-state information gain as a soft expansion potential and selects active prefixes with higher informativeness with higher probability, while gradually suppressing low-informativeness branches.\"},{\"question\":\"What benefit does the information gain-based rollout bring for policy optimization?\",\"answer\":\"It induces an explicit limiting teacher distribution over trajectories, which provides a clear policy optimization target and unifies adaptive tree-structured exploration with principled policy learning.\"}]",1784184928,48,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"information-gain-based-rollout-policy-optimization","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/information-gain-based-rollout-policy-optimization/83059/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-23","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does IGRPO address in long-horizon LLM search training?","Question",{"text":75,"@type":76},"It targets inefficient rollout budget allocation caused by not explicitly estimating how useful intermediate states are, leading to wasted computation on low-value or misleading branches.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does IGRPO decide which tree nodes to expand during rollouts?",{"text":80,"@type":76},"IGRPO treats intermediate-state information gain as a soft expansion potential and selects active prefixes with higher informativeness with higher probability, while gradually suppressing low-informativeness branches.",{"name":82,"@type":73,"acceptedAnswer":83},"What benefit does the information gain-based rollout bring for policy optimization?",{"text":84,"@type":76},"It induces an explicit limiting teacher distribution over trajectories, which provides a clear policy optimization target and unifies adaptive tree-structured exploration with principled policy learning.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},"General","general"]