[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-81610-en":3,"doc-seo-81610-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},81610,34359740700684,"Finn","https://ap-avatar.wpscdn.com/avatar/1f400023980c374ae676?_k=1777273430885731487",8,"Research & Report","SWE-Milestone：评估AI代理在持续软件演进中的表现","Real-world software must adapt to continuously emerging and open-ended requirements, making evolution an incremental, long-horizon process. Existing agent benchmarks test isolated, one-off coding problems and miss temporal dependencies and technical-debt accumulation. DeepCommit introduces an agentic pipeline reconstructing verifiable milestone-level milestone DAGs from noisy commit logs, enabling SWE-Milestone to assess agents on continuous streams. Experiments with 12 frontier models across 4 frameworks show a sharp drop from 80%+ on isolated tasks to at most 38% in continuous settings, highlighting weak long-term maintenance and error propagation.","arXiv :2603 . 13428v 3 [ cs . SE] 10 Jul 2026  \nSWE-Milestone: Evaluating AI Agentson Continuous Software Evolution  \nGangda Deng1,* , Zhaoling Chen2,* , Zhongming Yu3 , Haoyang Fan1 , Yuhong Liu1 , Yuxin Yang1 , Dhruv Parikh1 , Rajgopal Kannan4 , Le Cong5 , Mengdi Wang6 , Qian Zhang2 , Viktor Prasanna1 , Xiangru Tang7,†, Xingyao Wang8  \n1USC 2UCR 3UCSD 4Army Research Office 5 Stanford 6Princeton 7 Haven 8 OpenHands  \n* Equal Contribution †Corresponding Author  \n DeepCommit Pipeline: [github.com/DeepCommit-ai/DeepCommit](github.com/DeepCommit-ai/DeepCommit)  \n Benchmark: [github.com/DeepCommit-ai/SWE-Milestone](github.com/DeepCommit-ai/SWE-Milestone)  \n Data: [huggingface.co/datasets/DeepCommit-ai/SWE-Milestone-data](huggingface.co/datasets/DeepCommit-ai/SWE-Milestone-data)  \n Leaderboard: [swe-milestone.com](swe-milestone.com)  \nReal-world software must continuously evolve to meet ever-changing and open-ended requirements. AI agents, increasingly deployed as long-running systems, are now entrusted to drive this evolution. Yet, existing benchmarks evaluate agents on isolated, one-off coding tasks, neglecting the temporal dependencies and technical debt inherent in real-world software evolution. To bridge this gap, we introduce DeepCommit, an agentic pipeline that reconstructs verifiable Milestone DAGs from noisy commit logs, where milestones are defined as functionally cohesive development goals. These executable sequences enable SWE-Milestone, a benchmark that evaluates agents on streams of milestone-level tasks, requiring them to sustain system integrity and limit error accumulation, dimensions of long-term software evolution largely missing from current benchmarks. Our evaluation of 12 frontier models across 4 agent frameworks reveals a critical vulnerability: overall performance scores drop significantly from >80% on isolated tasks to at most 38% in continuous settings, exposing agents’ profound struggle with long-term maintenance and error propagation.  \nFigure 1 | Milestone-level task granularity optimally balances functional coherence and evolutionary awareness for benchmarking continuous software evolution.  \n1. Introduction  \nSoftware development in the real world is driven by dynamic, open-ended requirements. New requirements continuously emerge, with some building on earlier ones while others can be pursued in parallel. As a result, software evolves through an ongoing, incremental process rather than a one-time effort. Frontier LLM agents (e.g., Claude Code (Anthropic, 2025), Codex (OpenAI, 2025)) are increasingly entrusted to drive this evolution as long-running systems (Nous Research, 2026; OpenClaw, 2026), autonomously developing and refining software within complex environments rather than producing one-off edits. Over time, such an agent’s accumulated efforts naturally trace  \nContact: Gangda Deng ([gangdade@usc.edu](gangdade@usc.edu)), Zhaoling Chen ([zchen526@ucr.edu](zchen526@ucr.edu)), Xiangru Tang ([xiangru.tang@yale.edu](xiangru.tang@yale.edu))  \nTable 1 | Representative software engineering benchmarks for LLMs. Unlike other categories of benchmarks that evaluate agents on isolated snapshots or against ground-truth states at each step, Repository Evolution requires agents to continuously build upon their own accumulated development history, exposing them to error propagation across tasks. SWE-Milestone adopts the Milestonelevel granularity, a functionally coherent group of commits that collectively advance a development objective, avoiding the noise of individual commits and the excessive scope of full releases. Dev History denotes the agent’s own development trace accumulated from preceding tasks.  \n\n| Category | Benchmark | Language | Task Properties |  |  | Cross-task\u003Cbr>Dependency | Task\u003Cbr>Collection |\n| --- | --- | --- | --- | --- | --- | --- | --- |\n|  |  |  | Granularity | Avg LoC | Additional Context |  |  |\n| Function Completion | HumanEval (Chen et al., 2021) | Python | Function-level | 6.8 | –","cbCaioqAARplAan9","https://ap.wps.com/l/cbCaioqAARplAan9","pdf",4242555,1,44,"English","en",105,"# Introduction\n## Problem: isolated benchmarks ignore temporal evolution\n## DeepCommit and milestone DAG reconstruction\n## SWE-Milestone benchmark design\n## Experimental results and vulnerability finding","[{\"question\":\"Why do existing AI-agent benchmarks fail to reflect real software evolution?\",\"answer\":\"They evaluate agents on isolated, one-off coding tasks, neglecting temporal dependencies and the technical-debt accumulation that occurs during continuous development.\"},{\"question\":\"What does DeepCommit contribute to the SWE-Milestone benchmark?\",\"answer\":\"DeepCommit reconstructs verifiable milestone DAGs from noisy commit logs, defining milestones as functionally cohesive development goals.\"},{\"question\":\"What is the key evaluation finding across models and frameworks?\",\"answer\":\"Overall performance drops significantly in continuous settings, falling from above 80% on isolated tasks to at most 38%, showing major difficulties in long-term maintenance and error propagation.\"}]",1784174761,111,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"swe-milestone-evaluating-ai-agents-on-continuous-software-evolution","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/swe-milestone-evaluating-ai-agents-on-continuous-software-evolution/81610/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why do existing AI-agent benchmarks fail to reflect real software evolution?","Question",{"text":75,"@type":76},"They evaluate agents on isolated, one-off coding tasks, neglecting temporal dependencies and the technical-debt accumulation that occurs during continuous development.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What does DeepCommit contribute to the SWE-Milestone benchmark?",{"text":80,"@type":76},"DeepCommit reconstructs verifiable milestone DAGs from noisy commit logs, defining milestones as functionally cohesive development goals.",{"name":82,"@type":73,"acceptedAnswer":83},"What is the key evaluation finding across models and frameworks?",{"text":84,"@type":76},"Overall performance drops significantly in continuous settings, falling from above 80% on isolated tasks to at most 38%, showing major difficulties in long-term maintenance and error propagation.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]