[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-152230-en":3,"doc-seo-152230-105":30,"detail-sidebar-cat-0-en-105":84},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},152230,8796096645457,"Arica Lee","https://ap-avatar.wpscdn.com/avatar/800003749518d68ffe3?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779345340919836971",8,"Research & Report","Towards Self-Evolving Agent Benchmarks: Validatable Agent Trajectory via Test-Time Exploration","Recent advances in large language models (LLMs) and agent system designs have accelerated the capability ceiling of new agents, making existing benchmarks harder to use for reliable ability evaluation. The TRACE framework addresses this issue by evolving tasks dynamically from an existing benchmark, while recording execution trajectories. TRACE operates in three stages: proposal mining, free-exploration construction with trajectory capture, and multi-level validation to ensure reproducibility and logical coherence, improving complexity and correctness reliability on GAIA and adapting to reasoning benchmarks.","arXiv :2510 .00415v3 [ cs .AI] 24 Mar 2026  \nTOWARDS SELF-EVOLVING AGENT BENCHMARKS: VALIDATABLE AGENT TRAJECTORY VIA TEST-TIME EXPLORATION  \nDadi Guo 1 ,2∗, Tianyi Zhou 1 ,3∗, Dongrui Liu 1∗ , Chen Qian4 , Qihan Ren5 , Shuai Shao5 , Zhiyuan Fan,2 Yi R. Fung,2 Kun Wang,6 Linfeng Zhang,5 Jing Shao 1†  \n1 Shanghai Artificial Intelligence Laboratory 2Hong Kong University of Science and Technology  \n3University of Michigan 4Renmin University of China 5 Shanghai Jiao Tong University  \n6Nanyang Technological University [dguoae@connect.ust.hk](dguoae@connect.ust.hk) [zhtianyi@umich.edu](zhtianyi@umich.edu)[ ](zhtianyi@umich.edu){liudongrui, [shaojing}@pjlab.org.cn](shaojing}@pjlab.org.cn)  \nABSTRACT  \nRecent advances in large language models (LLMs) and agent system designs have empowered agents with unprecedented levels of capability. However, existing agent benchmarks are showing a trend of rapid ceiling-hitting by newly developed agents, making it increasingly difficult to meet the demands of evaluating agent abilities. To address this problem, we propose the Trajectory-based Validatedby-Reproducing Agent-benchmark Complexity Evolution (TRACE) framework.  \nThis framework takes an original task from an existing benchmark and encourages agents to freely explore and evolve it into a new task with higher difficulty while recording the corresponding execution trajectories. The framework proceeds in three stages: (1) evolutionary proposal mining, which generates task evolution proposals through preliminary exploration and divergent thinking; (2) problem construction via free exploration, where proposals are instantiated into concrete problem instances through agent exploration, with execution trajectories recorded along the process; and (3) multi-level validation, which ensures that the evolved tasks are accompanied by reproducible and logically coherent trajectories. Experiments on the GAIA benchmark demonstrate that the TRACE framework consistently enhances task complexity while improving correctness reliability through trajectory-level validation. In addition, our framework can successfully adapt to and improve reasoning benchmarks such as AIME-2024 . This work marks a paradigm shift from static, manually curated benchmarks to dynamic, self-evolving evaluation systems, providing a sustainable and challenging foundation for agent development.  \n1 INTRODUCTION  \nThe paradigm of artificial intelligence is rapidly shifting towards autonomous agents capable of complex reasoning (Comanici et al., 2025; Huang & Yang, 2025), planning (Huang et al., 2024), and tool utilization (Qu et al., 2025; Wang et al., 2024a) . This progress is starkly evident in the performance on challenging agent benchmarks (Mialon et al., 2023; Jimenez et al., 2023), which were once considered formidable. For instance, on the GAIA benchmark which is designed to test real-world assistant capabilities, top-performing agents have achieved scores exceeding 90%(GAIA Benchmark Team, 2025), rapidly closing the gap with the human baseline. This rapid pace signalsan urgent challenge: existing benchmarks are becoming saturated, diminishing their ability to differentiate state-of-the-art agents and risking progress being driven by overfitting to static test sets rather  \n∗Equal contribution. Work done during an internship at Shanghai Artificial Intelligence Laboratory, supervised by Dongrui Liu  \n†Corresponding author  \nPass@1 Score Across Different Levels and Rounds  \nFigure 1: Model performance comparison on the Pass@1 metric across four distinct difficulty levels and evolution rounds under the TRACE framework. As the number of evolution rounds increases, the performance of models shows a downward trend, demonstrating that our framework successfully evolves more challenging tasks.  \nthan generalizable intelligence. However, the cost of manually creating novel, complex, and reliable tasks is a labor-intensive, time-consuming, and expensive process, which highlights an urgent ne","cbCaiebT1DGhk03p","https://ap.wps.com/l/cbCaiebT1DGhk03p","pdf",3569686,1,31,"English","en",105,"# Abstract\n# 1 Introduction","[{\"question\":\"How does TRACE perform on the GAIA benchmark?\",\"answer\":\"Experiments on GAIA show that TRACE consistently increases task complexity while improving correctness reliability through trajectory-level validation, and it can also adapt to reasoning benchmarks such as AIME-2024.\"}]","Towards Self-Evolving Agent Benchmarks: Validatable Agent Trajectory via Test-Time Exploration | PDF",1787857926,78,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":79,"head_meta":81,"extra_data":83,"updated_unix":28},"towards-self-evolving-agent-benchmarks-validatable-agent-trajectory-via-test-time-exploration","",{"@graph":36,"@context":78},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/towards-self-evolving-agent-benchmarks-validatable-agent-trajectory-via-test-time-exploration/152230/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-09-06","2026-08-27",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72],{"name":73,"@type":74,"acceptedAnswer":75},"How does TRACE perform on the GAIA benchmark?","Question",{"text":76,"@type":77},"Experiments on GAIA show that TRACE consistently increases task complexity while improving correctness reliability through trajectory-level validation, and it can also adapt to reasoning benchmarks such as AIME-2024.","Answer","https://schema.org",{"og:url":52,"og:type":80,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":82,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":85},[86,90,94,98,103,108,113,116,121,124,128],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":87,"show_sort_weight":88,"slug":89},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":91,"show_sort_weight":92,"slug":93},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Exam",70,"exam",{"id":99,"doc_module":4,"doc_module_name":46,"category_name":100,"show_sort_weight":101,"slug":102},5,"Comic",60,"comic",{"id":104,"doc_module":4,"doc_module_name":46,"category_name":105,"show_sort_weight":106,"slug":107},6,"Technology",50,"technology",{"id":109,"doc_module":4,"doc_module_name":46,"category_name":110,"show_sort_weight":111,"slug":112},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":114,"slug":115},30,"research-report",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":118,"show_sort_weight":119,"slug":120},9,"Religion & Spirituality",20,"religion-spirituality",{"id":119,"doc_module":4,"doc_module_name":46,"category_name":122,"show_sort_weight":119,"slug":123},"World Cup","world-cup",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":126,"show_sort_weight":125,"slug":127},10,"Lifestyle","lifestyle",{"id":129,"doc_module":4,"doc_module_name":46,"category_name":130,"show_sort_weight":99,"slug":131},19,"General","general"]