[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82118-en":3,"doc-seo-82118-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},82118,1099514067415,"Rowan","https://ap-avatar.wpscdn.com/avatar/100002539d78ffe74a7?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779092875211072502",8,"Research & Report","Long-Horizon-Terminal-Bench：使用密集基于奖励的评分测试长任务中智能体的极限","Long-Horizon-Terminal-Bench is a challenging terminal benchmark focused on long-horizon, domain-specific agent workflows that existing terminal benchmarks largely miss. It introduces 46 tasks across nine categories, including experiment reproduction, software engineering, multimodal analysis, interactive games, and scientific computing. Each task is decomposed into fine-grained graded subtasks with intermediate checks to provide dense rewards and partial credit. Evaluation shows frontier agents require extensive execution time and tokens yet achieve low pass rates under partial and perfect reward thresholds.","arXiv :2607 .08964v 1 [ cs .AI] 9 Jul 2026  \nLong-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading  \nZongxia Li†1 ,2 , Zhongzhi Li†1 ,3 , Yucheng Shi†1 , Ruhan Wang 1 ,5 , Junyao Yang 1 ,7 , Zhichao Liu2 , Xiyang Wu2 , Anhao Li4 , Yue Yu5 , Ninghao Liu8 , Lichao Sun6 , Haotao Mi 1 , LeoweiLiang 1  \n1 Tencent HY LLM Frontier 2 University of Maryland, College Park 3 University of Georgia 4 University of Minnesota, Twin Cities 5 Indiana University 6 Lehigh University 7 National University of Singapore  \n8 The Hong Kong Polytechnic University † Core contribution.  \n[zli12321@umd.edu](zli12321@umd.edu) , [zongxiali@global.tencent.com](zongxiali@global.tencent.com) ,[zhongzhili@global.tencent.com](zhongzhili@global.tencent.com)  \nAI agents have become increasingly capable of autonomously completing short, well-specified tasks. However, existing terminal benchmarks largely focus on relatively simple problems that finish within a few minutes and are typically evaluated only by their final outcome. This setup overlooks intermediate progress and partial solutions, leading to sparse reward signals and an incomplete picture of agent capability.  \nWe introduce Long-Horizon-Terminal-Bench, a challenging terminal benchmark of 46 long-horizon tasks spanning nine categories, including experiment reproduction, software engineering, multimodal analysis, interactive games, and scientific computing. Each task follows a Terminal-Bench-style setup with a reference solution or simulation engine, but is further decomposed into fine-grained graded subtasks. This design enables dense intermediate rewards and partial credit, allowing evaluation to capture not only whether an agent reaches the final goal, but also how far it progresses on difficult, open-ended workflows.  \nTasks in Long-Horizon-Terminal-Bench typically require hundreds of episodes and tens of minutes to hours of execution, stressing long-horizon planning, long-context management, and iterative debugging rather than one-shot problem solving. We evaluate 15 frontier models and find that agents consume on average 9.9M tokens per task, with roughly 231 episodes and 85.3 minutes of execution time per run, making Long-Horizon-Terminal-Bench substantially more demanding than prior terminal-based benchmarks. Even the strongest tested model achieves 15.2% pass@1 at a partial-reward threshold of 0 .95 and 10.9% at a perfect-reward threshold of 1 .0, while the mean pass rate across models is just 4.3% and 1.7% under the two thresholds, respectively. These results reveal substantial headroom for improvement. We further analyze common failure modes and error patterns, and release Long-Horizon-Terminal-Bench to support future progress on robust long-horizon terminal agents.  \nProject page: [https://zli12321.github.io/LHTB/](https://zli12321.github.io/LHTB/)  \nDate: July 13, 2026  \n1 Introduction  \nAI agents built on large language models (LLMs) are rapidly improving at autonomous decision making and tool use [2, 6] . Recent work has demonstrated impressive performance on short, well-scoped tasks such as fixing a code repository issue, completing a coding ticket, or issuing a handful of shell commands [22] . However, they still cover only a narrow slice of the long-horizon, domain-specific, and practically important workflows that human experts care about in practice [22, 46 , 51] .  \nNowadays, many real workflows are long-horizon: they require agents to execute hundreds of steps, maintain and update plans over tens of minutes to hours, and manage evolving long-context state [42] . Examples include reproducing results from published research papers [36], installing environments from a github repo, auditing complex multimodal datasets, debugging compiler toolchains, or shepherding a multi-stage ML training pipeline. In these settings, agents must repeatedly use terminal command lines to read and write files, run scripts, inspect partial o","cbCaimroQimVLYq9","https://ap.wps.com/l/cbCaimroQimVLYq9","pdf",2636950,2,1,17,"English","en",105,"# Introduction\n## Motivation and gaps in existing benchmarks\n## Long-Horizon-Terminal-Bench design\n## Evaluation setup and results","[{\"question\":\"Long-Horizon-Terminal-Bench与以往终端基准相比，主要改进是什么？\",\"answer\":\"它将每个长任务分解为细粒度、带中间校验的分项子任务，从而实现密集的中间奖励与部分得分，而不仅仅看最终结果。\"},{\"question\":\"Long-Horizon-Terminal-Bench包含哪些任务类型和覆盖范围？\",\"answer\":\"共46个长时域终端任务，覆盖九类内容，包括实验复现、软件工程、多模态分析、交互游戏和科学计算等。\"},{\"question\":\"基准评测揭示了哪些性能现状与不足？\",\"answer\":\"评测显示模型平均需要大量tokens与较长执行时间；即使最强模型在部分奖励与完美奖励阈值下的pass@1也很低，不同模型的平均通过率整体偏低，体现了改进空间。\"}]",1784178298,43,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"long-horizon-terminal-bench-testing-the-limits-of-agents-on-long-horizon-terminal-tasks-with-dense-reward-based-grading","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/long-horizon-terminal-bench-testing-the-limits-of-agents-on-long-horizon-terminal-tasks-with-dense-reward-based-grading/82118/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-20","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Long-Horizon-Terminal-Bench与以往终端基准相比，主要改进是什么？","Question",{"text":75,"@type":76},"它将每个长任务分解为细粒度、带中间校验的分项子任务，从而实现密集的中间奖励与部分得分，而不仅仅看最终结果。","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Long-Horizon-Terminal-Bench包含哪些任务类型和覆盖范围？",{"text":80,"@type":76},"共46个长时域终端任务，覆盖九类内容，包括实验复现、软件工程、多模态分析、交互游戏和科学计算等。",{"name":82,"@type":73,"acceptedAnswer":83},"基准评测揭示了哪些性能现状与不足？",{"text":84,"@type":76},"评测显示模型平均需要大量tokens与较长执行时间；即使最强模型在部分奖励与完美奖励阈值下的pass@1也很低，不同模型的平均通过率整体偏低，体现了改进空间。","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]