[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84956-en":3,"doc-seo-84956-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},84956,7971461740886,"Theodore","https://ap-avatar.wpscdn.com/davatar_3d24733baf745e90a7e4bdd5f77d97b2",8,"Research & Report","MILES Modular Instruction Memory with Learnable Selection for Self-Improving LLM Reasoning","MILES (Modular Instruction Memory with LEarnable Selection) introduces a test-time self-improving framework for large language model reasoning that leverages sequential experience without updating model parameters. Instead of whole-solution templates or heuristic step selection, it maintains dynamically growing modular memory units as asymmetric (sub-goal embedding, sub-instruction) pairs. Learnable selection heads enable coarse-to-fine retrieval: a similarity stage expands memory and gathers supervision from confident trajectories, while a reranking stage optimizes correctness on uncertain cases. Experiments show improved accuracy–efficiency tradeoffs, robustness, and transferability.","arXiv :2607 .06974v 1 [ cs .CL] 8 Jul 2026  \nMILES: Modular Instruction Memory with Learnable Selection for Self-Improving LLM Reasoning  \nRuilin Tong , Dong GongB  \nUniversity of New South Wales (UNSW Sydney)  \n{ruilin.tong, [dong.gong](dong.gong}@unsw.edu.au)[}](dong.gong}@unsw.edu.au)[@unsw.edu.au](dong.gong}@unsw.edu.au)  \n[B](BCorresponding author.)[Corresponding author.](BCorresponding author.)  \nLarge language models (LLMs) increasingly improve their reasoning at test time via additional computation, yet most existing works treat each problem in isolation. When problems arrive sequentially, accumulating reusable experience across them can further improve performance. Existing memory-based methods either store whole-solution templates that generalize poorly to novel problems or use heuristic step-level selection that is not optimized for final-answer correctness. Learning selection policies requires large-scale training data and fixed action spaces, making such approaches unsuitable for test-time settings where memory expands incrementally and only limited supervision is available. We propose MILES (Modular Instruction Memory with LEarnable Selection for self-improving LLM reasoning), a framework that dynamically expands step-wise memory and applies correctness-optimized memory composition under realistic test-time constraints. MILES maintains modular memory units consisting of asymmetric pairs of sub-goal embeddings and sub-instructions, each associated with a learnable selection head. This memory structure enables a coarse-to-fine retrieval mechanism: The coarse level enables memory expansion and collects supervision for training selection heads from confident samples, while the fine stage applies learned selection heads to rerank coarse-level candidates and guide reasoning for uncertain samples. MILES consistently matches or outperforms prior methods while achieving superior accuracy–efficiency tradeoffs.  \nExtensive experiments demonstrate its effectiveness, robustness, and transferability.  \nProject Page: MILES  \n1 Introduction  \nLarge language models (LLMs) demonstrate strong capabilities in complex reasoning and question answering. Recent advances in test-time computation, including chain-of-thought prompting [42], self-consistency [40], and tree search [48, 28], have further improved reasoning performance. In many practical settings, however, problems arrive sequentially rather than in isolation, allowing models to benefit from accumulating reusable experience across related queries [47, 26] . Updating model parameters for every new problem is often costly or infeasible, especially for closed or frozen models [50, 51] . A natural alternative is therefore to maintain an external memory that the LLM consults at test time while its parameters remain frozen [35, 53 , 33] . The central question becomes: how can such a memory be organized, maintained, and reused so that it actually improves reasoning—and continues to improve as more experience accumulates?  \nExisting memory methods for LLM reasoning store whole-solution templates or strategies and retrieve them per problem [47, 35 , 53], which are useful when the new problem closely resembles a seen one, but limited when the new problem has novel structure. Step-level memory [3, 34 , 12] stores smaller units that can be composed flexibly across reasoning steps and across problems with different overall structures, shifting the problem from finding a similar past problem to finding applicable local blocks for the next reasoning step. Among step-level designs, prior work commits to all-text concept/situation pairs [12], pre-formed templates [3], or free-form strategies [33, 26] . A common limitation persists across all of these: selecting which unit to apply at the current step is heuristic, via similarity retrieval or LLM prompting, without optimizing the decision against final-answer correctness. Methods that learn selection from feedback [55, 6 , 56] require fixed act","cbCaivqlTgjpjafJ","https://ap.wps.com/l/cbCaivqlTgjpjafJ","pdf",1989963,2,1,34,"English","en",105,"# Introduction\n## Problem Setting: Sequential Queries and Frozen LLMs\n## Limitations of Existing Memory and Selection Methods\n## Proposed Framework: MILES and Learnable Memory Selection","[{\"question\":\"Why do existing memory-based LLM reasoning methods underperform on novel problem structures?\",\"answer\":\"They often store whole-solution templates or use step-level units but rely on heuristic selection that is not optimized for final-answer correctness, limiting generalization when structures differ.\"},{\"question\":\"How does MILES structure its memory units?\",\"answer\":\"MILES stores modular asymmetric pairs of a sub-goal embedding (retrieval key) and a natural-language sub-instruction (generation guidance), enabling reusable local reasoning modules.\"},{\"question\":\"How does MILES learn and apply selection during test time?\",\"answer\":\"A coarse-to-fine mechanism is used: a similarity-based retrieval stage expands memory and collects supervision from confident samples to train selection heads, then a fine stage reranks candidates and guides reasoning for uncertain samples for correctness-optimized composition.\"}]",1784199688,86,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"miles-modular-instruction-memory-with-learnable-selection-for-self-improving-llm-reasoning","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/miles-modular-instruction-memory-with-learnable-selection-for-self-improving-llm-reasoning/84956/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-23","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why do existing memory-based LLM reasoning methods underperform on novel problem structures?","Question",{"text":75,"@type":76},"They often store whole-solution templates or use step-level units but rely on heuristic selection that is not optimized for final-answer correctness, limiting generalization when structures differ.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does MILES structure its memory units?",{"text":80,"@type":76},"MILES stores modular asymmetric pairs of a sub-goal embedding (retrieval key) and a natural-language sub-instruction (generation guidance), enabling reusable local reasoning modules.",{"name":82,"@type":73,"acceptedAnswer":83},"How does MILES learn and apply selection during test time?",{"text":84,"@type":76},"A coarse-to-fine mechanism is used: a similarity-based retrieval stage expands memory and collects supervision from confident samples to train selection heads, then a fine stage reranks candidates and guides reasoning for uncertain samples for correctness-optimized composition.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]