[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84257-en":3,"doc-seo-84257-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},84257,13056703019662,"Evangeline","https://ap-avatar.wpscdn.com/avatar/be000253a8e92610077?_k=1778726343310543188",8,"Research & Report","Recursive Self-Improvement in AI From Bounded Self-Refinement to Autonomous Research Loops","AI systems increasingly improve themselves by revising outputs, adapting toolchains during deployment, training on self-generated data, and—within a fast-growing line of work—conducting AI research. A survey of 1,250 arXiv papers (2024–2026) disambiguates a conflated “self-X” vocabulary by organizing methods along what is improved (deployment behavior, training policy, evaluators, or the research process) and how closed the loop is. It distinguishes bounded self-refinement from open-ended RSI, analyzes evaluator-design space, failure modes, and governance-grade measurement gaps.","arXiv :2607 .07663v 1 [ cs .AI] 8 Jul 2026  \nRecursive Self-Improvement in AI: From Bounded Self-Refinement  \nto Autonomous Research Loops  \nMingguang Chen 1,∗ Licheng Wang2 Bo Qu3  \n1 University of California, Riverside (UCR) 2 AlphaAvatar 3 Illinois Institute of Technology (IIT)  \n∗ Corresponding [authors. Email: mchen041@ucr.edu](authors. Email: mchen041@ucr.edu)  \nJuly 2026  \nAbstract  \nAI systems increasingly participate in their own improvement: revising their outputs, adapting and evolving their own harnesses during deployment, training on data they generate, and  \n— in a growing research thread — conducting AI research itself. The literature describing this participation has exploded, but under a vocabulary (“self-refine,” “self-reward,” “self-play,”“self-evolve”) that conflates fundamentally different ambitions. We survey 1,250 arXiv papers (2024–2026) and organize them along two axes: what the system improves — its behavior in deployment, its policy through training, its evaluator, or the research process itself—and the degree of loop closure (human-in-the-loop to fully closed) . The taxonomy separates bounded self-refinement — convergent, evaluable, and already industrial practice — from open-ended recursive self-improvement (RSI), which remains bounded by grounding requirements, collapse dynamics, and compute constraints on every side current evidence can measure. Its distinctive feature is a dedicated category for self-evaluation: every improvement loop is a claim that some signal can substitute for human judgment, so we survey the evaluator design space —judges, process reward models, verifiers, rubrics, meta-evaluation — alongside the loops it supervises, order the signals into a verification hierarchy from formal verifiers (strongest) to intrinsic selfassessment (weakest), and observe that demonstrated self-improvement strength tracks this hierarchy, that its characteristic failure modes (self-confirming loops, model collapse, diversity collapse) follow from its violations, and that the “research direction-setting” bottleneck that keeps humans in the loop is its top rung. We connect the technical literature to the theory of RSI limits and to the safety and governance questions raised by frontier-lab accounts of closing the loop, and identify governance-grade measurement of self-improvement as the field’s most underpopulated niche.  \n1 . Introduction  \nThe idea that an artificial intelligence might improve itself—and that each improvement might make the next one easier—is among the oldest in the field. Good’s “intelligence explosion” argument [1] and Schmidhuber’s provably-optimal Gödel machines [2] framed recursive self-improvement (RSI) as a theoretical endpoint decades before any system could plausibly attempt it. What has changed is that fragments of the loop are now engineering practice. Large language models routinely critique and revise their own outputs, train on data they themselves generated, rewrite their own agent scaffolding, and—in systems like FunSearch [3] and AlphaEvolve [4]—discover algorithms that feed back into the infrastructure of AI development itself.  \nAnthropic’s recent essay on recursive self-improvement [5] frames this transition as a continuum of increasing AI autonomy in the AI-improvement loop: from humans writing all code (pre-2023), through chatbot-assisted coding and autonomous coding agents, to agents that delegate work to other agents today, and—at the end of the spectrum—“closing the loop”: agents that design and train their successor models. The essay argues that current systems sit conspicuously far along this spectrum on execution (as of May 2026, Claude reportedly writes over 80% of Anthropic’s merged code) while remaining bottlenecked on research direction-setting—choosing which problems matter. Whether, when, and how the remaining gap closes is arguably the most consequential open question in the field. We use the essay as a motivating frame and a source of stage vocabu","cbCaiqaPC3jYMAmb","https://ap.wps.com/l/cbCaiqaPC3jYMAmb","pdf",5035755,6,1,42,"English","en",105,"# Abstract\n# Introduction\n## Motivation and survey scope\n## Taxonomy axes: what is improved and loop closure\n## Acceleration vs. consolidation","[{\"question\":\"What does the survey focus on regarding AI self-improvement?\",\"answer\":\"It surveys 1,250 arXiv papers (2024–2026) and organizes them by what the system improves (deployment behavior, training policy, evaluator, or research process) and by the degree of loop closure (human-in-the-loop to fully closed).\"},{\"question\":\"How does the paper distinguish bounded self-refinement from open-ended recursive self-improvement (RSI)?\",\"answer\":\"Bounded self-refinement is convergent and evaluable, aligning with already industrial practice, while open-ended RSI remains limited by grounding requirements, collapse dynamics, and compute constraints. The paper highlights self-evaluation as a core category within the improvement loop.\"},{\"question\":\"Why does the survey emphasize evaluator design in self-improvement loops?\",\"answer\":\"Each improvement loop implicitly claims some signal can substitute for human judgment, so it surveys evaluator designs (judges, process reward models, verifiers, rubrics, meta-evaluation) and links stronger demonstrated self-improvement to higher levels in a verification hierarchy.\"}]",1784194413,106,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"recursive-self-improvement-in-ai-from-bounded-self-refinement-to-autonomous-research-loops","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/recursive-self-improvement-in-ai-from-bounded-self-refinement-to-autonomous-research-loops/84257/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What does the survey focus on regarding AI self-improvement?","Question",{"text":76,"@type":77},"It surveys 1,250 arXiv papers (2024–2026) and organizes them by what the system improves (deployment behavior, training policy, evaluator, or research process) and by the degree of loop closure (human-in-the-loop to fully closed).","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does the paper distinguish bounded self-refinement from open-ended recursive self-improvement (RSI)?",{"text":81,"@type":77},"Bounded self-refinement is convergent and evaluable, aligning with already industrial practice, while open-ended RSI remains limited by grounding requirements, collapse dynamics, and compute constraints. The paper highlights self-evaluation as a core category within the improvement loop.",{"name":83,"@type":74,"acceptedAnswer":84},"Why does the survey emphasize evaluator design in self-improvement loops?",{"text":85,"@type":77},"Each improvement loop implicitly claims some signal can substitute for human judgment, so it surveys evaluator designs (judges, process reward models, verifiers, rubrics, meta-evaluation) and links stronger demonstrated self-improvement to higher levels in a verification hierarchy.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,111,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":107,"slug":138},19,"General","general"]