[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83327-en":3,"doc-seo-83327-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83327,1374391974585,"Genevieve","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","TTHE: Test-Time Harness Evolution","The behavior of an LLM agent depends not only on the underlying model but also on its harness: a control program that builds context, calls tools, checks intermediate results, and recovers from failures. Most methods optimize a fixed harness before deployment, limiting adaptation when test conditions differ. Test-Time Harness Evolution (TTHE) optimizes the executable harness during evaluation using only unlabeled execution traces, maintaining candidate harness populations, judging via execution-derived proxy signals, and persisting improved programs. TTHE keeps LLM weights frozen and requires no gold labels or separate adaptation model.","arXiv :2607 .08 124v 1 [ cs . SE] 9 Jul 2026  \nTTHE: TEST-TIME HARNESS EVOLUTION  \nJun Nie 1 ,2 , Yonggang Zhang3 , Jun Song 1 , Qianshu Cai2 , Dahai Yu4 ,  \nYike Guo3 , Xinmei Tian2 ,∗ , Bo Han 1 ,∗  \n1Hong Kong Baptist University 2University of Science and Technology of China  \n3Hong Kong Generative AI Research & Development Center, The Hong Kong University of Science and Technology  \n4TCL Corporate Research (HK) Co., Ltd  \nABSTRACT  \nThe behavior of an LLM agent is determined not only by the underlying model, but also by its harness: the executable program that constructs context, invokes tools, verifies intermediate results, and recovers from failures. Existing approaches optimize such harnesses before deployment, searching training or development data for a fixed agent workflow that is then frozen at test time. This limits adaptation when the test distribution, failure modes, or tool interactions differ from those seen during development. We ask whether the harness can instead be optimized during evaluation itself, using only the unlabeled execution traces the agent produces on the test inputs. We introduce Test-Time Harness Evolution (TTHE), which treats the executable harness as the state of test-time adaptation. During evaluation, TTHE maintains a population of candidate harnesses and refines them through an agentic proposer that reasons over their execution traces, without gold labels or task-specific supervision; a judge then commits an improved harness from execution-derived proxy signals, and the selected program persists to govern subsequent inputs. Crucially, TTHE does not update model weights, require gold labels, or train a separate adaptation model: solver, proposers, and judge are different roles and harnesses around the same frozen LLM, so all adaptation occurs through changes to the surrounding program. Across text-to-SQL, competitive programming, software engineering, data-science coding, and agentic tool-use tasks, TTHE improves fixed ReAct-style baseline harnesses, yielding persistent, inspectable improvements rather than a pre-searched workflow or per-query retries. These results recast test-time adaptation for LLM agents as evolution over executable control programs and identify execution-derived proxy reliability as a central challenge for robust unsupervised agent improvement.  \n1 INTRODUCTION  \nModern LLM agents are increasingly built as executable systems rather than single-call predictors: they retrieve context, call tools, execute code, inspect intermediate states, and recover from failures. In such systems, the agent’s behavior is governed not only by the frozen model, but also by the executable harness around it—the control program that constructs context, invokes tools, verifies intermediate results, and implements recovery strategies. Following prior work on model and agent harnesses (Lee et al., 2026; Zhang et al., 2026), we focus on this harness as the object of adaptation. This surrounding program can shape an agent’s behavior as much as the model itself: the same model, when placed inside different harnesses, can produce substantially different outcomes on the same workload. In most deployed or benchmarked agent systems, however, the harness is fixed during evaluation. Practitioners inspect development failures, adjust prompts and control flow, tune tool-use logic, and then ship a static artifact whose behavior remains unchanged at test time.  \nThis convention treats evaluation as a passive scoring stage, even though evaluation itself produces rich operational evidence. Every attempted task leaves an execution trace: model calls, tool actions,  \n∗ Corresponding authors.  \nCode: [https://github.com/junnie00/TTHE](https://github.com/junnie00/TTHE)  \nintermediate outputs, database responses, test results, runtime errors, and recovery decisions. Such traces often expose not only whether an agent succeeded under available proxies, but also how its operating procedure might be improved. A text-","cbCaiuTAnjyI7McJ","https://ap.wps.com/l/cbCaiuTAnjyI7McJ","pdf",455013,3,1,15,"English","en",105,"# Abstract\n# Introduction\n# Related Work\n# Method: Test-Time Harness Evolution (TTHE)","[{\"question\":\"What problem does TTHE address for LLM agents?\",\"answer\":\"It addresses the limitation of fixed harnesses during evaluation, where the agent cannot adapt to changed test distributions, failure modes, or tool interactions. TTHE seeks harness improvements using only evidence produced during evaluation.\"},{\"question\":\"How does TTHE optimize the harness during evaluation?\",\"answer\":\"TTHE maintains a population of candidate executable harnesses and refines them with an agentic proposer that reasons over execution traces. A judge then selects an improved harness using execution-derived proxy signals, and the selected program persists for later inputs.\"},{\"question\":\"Does TTHE update the LLM model weights or require gold labels?\",\"answer\":\"No. TTHE keeps the LLM weights frozen, does not use gold labels, and does not train a separate adaptation model. All adaptation happens through changes to the surrounding executable harness.\"}]",1784186752,38,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"tthe-test-time-harness-evolution","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/tthe-test-time-harness-evolution/83327/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does TTHE address for LLM agents?","Question",{"text":75,"@type":76},"It addresses the limitation of fixed harnesses during evaluation, where the agent cannot adapt to changed test distributions, failure modes, or tool interactions. TTHE seeks harness improvements using only evidence produced during evaluation.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does TTHE optimize the harness during evaluation?",{"text":80,"@type":76},"TTHE maintains a population of candidate executable harnesses and refines them with an agentic proposer that reasons over execution traces. A judge then selects an improved harness using execution-derived proxy signals, and the selected program persists for later inputs.",{"name":82,"@type":73,"acceptedAnswer":83},"Does TTHE update the LLM model weights or require gold labels?",{"text":84,"@type":76},"No. TTHE keeps the LLM weights frozen, does not use gold labels, and does not train a separate adaptation model. All adaptation happens through changes to the surrounding executable harness.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]