[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82973-en":3,"doc-seo-82973-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},82973,687197207639,"Asher","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","What Do AI Agents Actually Change An Empirical Taxonomy of Mutation Patterns in Performance Improving Pull Requests","AI coding agents behave as black boxes, but their committed code diffs reveal how they transform software. This paper uses search-based software engineering and genetic improvement to motivate an empirical need for mutation operators grounded in observed agent edits. From 33,596 agent PRs, fewer than 1% target performance. The study analyzes 1,254 performance-relevant diff hunks across five agent systems, classifying them against an 18-category syntactic mutation taxonomy using a dual-LLM pipeline.","arXiv :2607 .05666v 1 [ cs . SE] 6 Jul 2026  \nWhat Do AI Agents Actually Change? An Empirical Taxonomy of Mutation Patterns in Performance-Improving Pull Requests  \nIllia Dovhoshliubnyi 1 , Nima Soroush 1 , Ashkan Sami 1 , and Alexander Brownlee2  \n1 Edinburgh Napier University, Edinburgh, UK  \n[illiadovho@gmail.com](illiadovho@gmail.com) , {N.Soroush, [A.Sami](A.Sami}@napier.ac.uk)[}](A.Sami}@napier.ac.uk)[@napier.ac.uk](A.Sami}@napier.ac.uk)  \n2 University of Stirling, Stirling, UK  \n[alexander.brownlee@stir.ac.uk](alexander.brownlee@stir.ac.uk)  \nAbstract. AI coding agents are black boxes: we cannot inspect how they  \ngenerate code, but we can inspect what they change. This distinction  \nmatters for search-based software engineering (SBSE), where techniques  \nsuch as genetic improvement (in the performance-optimisation application  \nwe study) depend on mutation operators that reflect how code is actually  \ntransformed. Fewer than 1% of the 33,596 agent PRs in AIDev-pop target  \nperformance, making each case a rare window into otherwise opaque agent  \nbehaviour. We classify 1,254 performance-relevant diff hunks from 216 of  \nthese PRs, spanning five agent systems, against the 18-category syntactic mutation taxonomy of Even-Mendoza et al. (2025) using a dual-LLM intersection pipeline. Three categories dominate: name   modification (37.0%), object   creation (26.4%), and type   change (22.7%), a profile markedly different from prior GI corpora where no   change accounted for 84% . Each  \nagent’s deployed system commits to a distinctive mutation vocabulary, and each performance strategy activates a largely disjoint category subset.  \nAgent identity and target strategy are therefore informative priors that  \nnarrow the effective SBSE operator space.  \nReplication package: [https://github.com/5uper6rain/ssbse-challenge-2026](https://github.com/5uper6rain/ssbse-challenge-2026)  \nKeywords: mutation testing · AI agents · empirical study · search-based software engineering · performance optimization  \n1 Introduction  \nAI coding agents such as Devin, GitHub Copilot, Cursor, OpenAI Codex, and Claude Code autonomously submit pull requests to production repositories, but their mechanisms for deciding what to change are opaque. We cannot inspect those mechanisms, but we can inspect the outputs: the actual code transformations they commit.  \nThis distinction matters for SBSE. Genetic improvement and related techniques depend on mutation operators grounded in empirical evidence of how code is actually transformed [2,3 ,7] . We focus specifically on the performanceimprovement application of GI. Classical mutation taxonomies were derived  \n2 I. Dovhoshliubnyi et al.  \nfrom human-written patches. As SBSE is increasingly applied to agent-assisted workflows, an empirical map of how agents actually transform code is a missing prior: descriptive of agent behaviour for the purpose of scoping operator selection in agent-aware tooling, not prescriptive of agent behaviour as an optimum.  \nPerformance-improving PRs are rare: only 324 of the 33,596 PRs in AIDevpop [4] carry a performance label (\u003C1%) [5,6], as most agent PRs target bugs or features. These rare cases are the instances where agents intentionally optimised code, making post-hoc mutation analysis directly informative for SBSE.  \nWe address two research questions:  \nRQ1 . What syntactic mutation patterns characterise successful performance PRs from AI coding agents, and how do they differ from those observed in prior genetic improvement work?  \nRQ2 . Do mutation patterns vary systematically across agent systems and across performance strategies?  \n2 Dataset and Methodology  \n2.1 Dataset  \nWe use the AIDev-pop subset of the AIDev dataset [4]: PRs from five AI coding agents (Devin, GitHub Copilot, Cursor, OpenAI Codex, and Claude Code) against 100 starred repositories. Of the 324 PRs carrying a performance label [6], 280 had a retrievable diff; 269 contained at least one source-code hunk after ","cbCaidlHPD7hPxw3","https://ap.wps.com/l/cbCaidlHPD7hPxw3","pdf",372130,2,1,6,"English","en",105,"# Abstract\n# Introduction\n# Dataset and Methodology\n## Dataset\n## Mutation Taxonomy\n# Mutation Patterns in AI-Agent Performance PRs","[{\"question\":\"Why is inspecting AI agent behavior challenging, and what does this paper inspect instead?\",\"answer\":\"AI coding agents are treated as black boxes, so their code generation mechanisms are not directly inspectable. The paper instead analyzes the committed code transformations in pull-request diffs.\"},{\"question\":\"What data and filtering steps are used to collect performance-relevant cases?\",\"answer\":\"The study uses AIDev-pop PRs from five agents and starts from 324 performance-labeled PRs. Diffs are decomposed into change hunks, then a dual-LLM intersection filter retains only performance-relevant hunks and removes cosmetic-only edits and false positives.\"},{\"question\":\"Which mutation categories dominate in performance-improving agent pull requests?\",\"answer\":\"The largest categories are name modification (37.0%), object creation (26.4%), and type change (22.7%). Together they define a performance-edit profile that differs from prior genetic improvement corpora.\"}]",1784184407,15,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"what-do-ai-agents-actually-change-an-empirical-taxonomy-of-mutation-patterns-in-performance-improving-pull-requests","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/what-do-ai-agents-actually-change-an-empirical-taxonomy-of-mutation-patterns-in-performance-improving-pull-requests/82973/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why is inspecting AI agent behavior challenging, and what does this paper inspect instead?","Question",{"text":75,"@type":76},"AI coding agents are treated as black boxes, so their code generation mechanisms are not directly inspectable. The paper instead analyzes the committed code transformations in pull-request diffs.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What data and filtering steps are used to collect performance-relevant cases?",{"text":80,"@type":76},"The study uses AIDev-pop PRs from five agents and starts from 324 performance-labeled PRs. Diffs are decomposed into change hunks, then a dual-LLM intersection filter retains only performance-relevant hunks and removes cosmetic-only edits and false positives.",{"name":82,"@type":73,"acceptedAnswer":83},"Which mutation categories dominate in performance-improving agent pull requests?",{"text":84,"@type":76},"The largest categories are name modification (37.0%), object creation (26.4%), and type change (22.7%). Together they define a performance-edit profile that differs from prior genetic improvement corpora.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,114,119,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":111,"show_sort_weight":112,"slug":113},"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]