[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84805-en":3,"doc-seo-84805-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},84805,2336464648322,"Aria","https://ap-avatar.wpscdn.com/avatar/2200025388227c56fec?_k=1778556882303663488",8,"Research & Report","MetaSkill-Evolve Recursive Self-Improvement of LLM Agents via Two-Timescale Meta-Skill Evolution","MetaSkill-Evolve presents a two-timescale framework for recursive self-improvement in LLM agents that handle long-horizon, open-ended tasks and rely on reusable skill specifications. Existing self-improving agents update only task skills while keeping the improvement procedure fixed, limiting adaptation when repeated failures share an unchanged diagnosis style. MetaSkill-Evolve co-evolves a branch-local meta-skill alongside evolving task skills, using the same improvement pipeline without extra models or objectives. Results on OfficeQA, SealQA, and ALFWorld improve held-out accuracy by +23.54, +16.09, and +1.92 points respectively.","MetaSkill-Evolve: Recursive Self-Improvement of LLM Agents via Two-Timescale Meta-Skill Evolution  \nZefeng Wang*,1 , Minxi Yan*,2 , Jinhe Bi1 , Sikuan Yan1 , Volker Tresp1 , Yunpu Ma1,3,4  \n1LMU Munich, 2The Chinese University of Hong Kong, 3MCML, 4MemAgents Lab  \narXiv :2607 .05297v 1 [ cs .AI] 6 Jul 2026  \nAbstract  \nRecent LLM agents tackle increasingly longhorizon, open-ended tasks, and external skills, reusable procedural knowledge supplied to the agent, further extend this capability. However, a fixed, hand-authored skill is rarely optimal, and cannot adapt to the diversity of tasks an agent encounters. Self-improving agents address this by rewriting their own skill files from execution traces, yielding meaningful gains on challenging benchmarks. Yet such self-evolution remains non-recursive: it improves only the task skill (what the agent does) while the improvement procedure (how it improves) is authored once and held fixed. We introduce MetaSkill-Evolve, a two-timescale framework that makes agentic skill improvement recursive: every branch carries both a task skill s and a branch-local meta-skill m = (ψ,σ,α,π,ε) whose five components parameterise the Analyzer, Retriever, Allocator, Proposer, and Evolver agents of the improvement pipeline. Task skills evolve on a fast loop while the meta-skill evolves on aslower one under the same pipeline applied to itself, with no additional model or objective. With all five pipeline agents sharing a single frozen backbone, MetaSkill-Evolve outperforms no-skill, static-skill, and single-level evolution baselines on three agentic benchmarks (OfficeQA, SealQA, ALFWorld), improving held-out test accuracy over the raw backbone by +23.54, +16.09, and +1.92 points respectively.  \n1 Introduction  \nLanguage model agents now tackle increasingly long-horizon, open-ended tasks, from document understanding and multi-step reasoning to tool use, yet they rarely succeed out of the box (Yao et al., 2023) . A productive remedy is to equip the agent  \n*  \nEqual contribution.  \nwith a skill: a curated, editable Markdown specification of reusable procedures, now a portable file-system artifact in widely deployed agent harnesses (Wang et al., 2023 ; Zheng et al., 2025) . But a fixed, hand-authored skill is rarely optimal, and cannot anticipate the diversity of tasks an agent encounters. Self-improvement systems such as EvoSkill (Alzubi et al., 2026), GEPA (Agrawalet al., 2026), and SkillWeaver (Zheng et al., 2025) address this by closing the loop with an analyze– propose–evolve pipeline that rewrites the skill after each failure trace, so that iteration by iteration the skill grows more capable.  \nThese systems, however, evolve only what the agent does, not how it evolves: the artifact under optimization changes while the operator that optimizes it stays fixed. In the vocabulary of selfimproving machines (Good, 1965 ; Schmidhuber, 2006), they are self-improving but stop short of being recursively self-improving. The meta-level logic is hardcoded in advance and shared by every branch throughout the run: how failures are diagnosed, which edits are proposed, how much search effort is allocated, whether cross-branch experience is reused, and how an approved edit is applied to disk (Fig. 1, third panel) . A branch therefore cannot improve the way it diagnoses failures: it applies the same procedure to every error, whether a misread table or a faulty calculation, and when that procedure yields the wrong fix, nothing in the loop can revise it.  \nA closer look at this rigidity suggests that two quantities govern evolutionary skill search. The first is the current skill utility U (s), the score of the present skill on a validation batch. The second is the meta-productivity P (m | s), the rate at which a branch generates stronger descendants under its current improvement policy m. These are not the same: a skill may score well today yet sit ina branch whose meta-level policy produces weak children, while a moderat","cbCaigeuBO9yE8ST","https://ap.wps.com/l/cbCaigeuBO9yE8ST","pdf",2117514,1,14,"English","en",105,"# Abstract\n# Introduction","[{\"question\":\"What problem does MetaSkill-Evolve address in current self-improving LLM agents?\",\"answer\":\"It addresses the limitation that most systems improve only the task skill while leaving the improvement procedure fixed, so the agent cannot revise how it diagnoses failures when those failures follow the same rigid style.\"},{\"question\":\"How does MetaSkill-Evolve make agent skill improvement recursive?\",\"answer\":\"Each branch maintains both a task skill s and a branch-level meta-skill m=(ψ,σ,α,π,ε). Task skills evolve on a fast loop, while the meta-skill evolves on a slower loop under the same pipeline applied to itself.\"},{\"question\":\"What performance gains are reported for MetaSkill-Evolve?\",\"answer\":\"On OfficeQA, SealQA, and ALFWorld, it improves held-out test accuracy over the raw backbone by +23.54, +16.09, and +1.92 points, outperforming no-skill, static-skill, and single-level evolution baselines.\"}]",1784198367,35,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"metaskill-evolve-recursive-self-improvement-of-llm-agents-via-two-timescale-meta-skill-evolution","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/metaskill-evolve-recursive-self-improvement-of-llm-agents-via-two-timescale-meta-skill-evolution/84805/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does MetaSkill-Evolve address in current self-improving LLM agents?","Question",{"text":75,"@type":76},"It addresses the limitation that most systems improve only the task skill while leaving the improvement procedure fixed, so the agent cannot revise how it diagnoses failures when those failures follow the same rigid style.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does MetaSkill-Evolve make agent skill improvement recursive?",{"text":80,"@type":76},"Each branch maintains both a task skill s and a branch-level meta-skill m=(ψ,σ,α,π,ε). Task skills evolve on a fast loop, while the meta-skill evolves on a slower loop under the same pipeline applied to itself.",{"name":82,"@type":73,"acceptedAnswer":83},"What performance gains are reported for MetaSkill-Evolve?",{"text":84,"@type":76},"On OfficeQA, SealQA, and ALFWorld, it improves held-out test accuracy over the raw backbone by +23.54, +16.09, and +1.92 points, outperforming no-skill, static-skill, and single-level evolution baselines.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]