[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83269-en":3,"doc-seo-83269-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83269,13056703019662,"Evangeline","https://ap-avatar.wpscdn.com/avatar/be000253a8e92610077?_k=1778726343310543188",8,"Research & Report","What Makes a Good Bug Report for an AI Agent","Automated program repair (APR) agents are moving from research benchmarks into real developer workflows, yet they still rely on bug reports originally written for humans. The study evaluates which bug-report features transfer to LLM-based agents using two complementary analyses. Statistical modeling links 27 report features with repair success across 433 SWE-bench Verified issues, highlighting the value of executable and well-localized information while showing longer reports may hurt. Controlled ablations further reveal strong dependence on localization cues and expected behavior, with differing handling of missing information between Qwen and Gemma.","What Makes a Good Bug Report for an AI Agent?  \nLara Khatib  \nUniversity of Waterloo Waterloo, Canada [lara.khatib@uwaterloo.ca](lara.khatib@uwaterloo.ca)  \nNoble Saji Mathews  \nUniversity of Waterloo Waterloo, Canada [noblesaji.mathews@uwaterloo.ca](noblesaji.mathews@uwaterloo.ca)  \nMeiyappan Nagappan  \nUniversity of Waterloo Waterloo, Canada [mei.nagappan@uwaterloo.ca](mei.nagappan@uwaterloo.ca)  \nPengyu Nie University of Waterloo Waterloo, Canada [pynie@uwaterloo.ca](pynie@uwaterloo.ca)  \nThomas Zimmermann  \nUniversity of California, Irvine Irvine, USA[tzimmer@uci.edu](tzimmer@uci.edu)  \narXiv :2607 .07593v 1 [ cs . SE] 8 Jul 2026  \nAbstract—Automated program repair (APR) agents are transitioning from research benchmarks to developer workflows, yet they still begin with bug reports written for human developers. While decades of research have established what makes a good bug report for humans (e.g., steps to reproduce, stack traces), it remains unclear whether these features transfer to LLM-based agents. We study this question in two complementary analyses. First, we use statistical modeling to examine associations between 27 bug-report features and repair success across 433 SWEbench Verified issues attempted by 87 repair agents. We find that fix suggestions, reproduction scripts, repository source code, and localization info are each associated with higher resolution likelihood, while longer reports are associated with lower odds. Second, we conduct controlled ablations across 2 models and 17 problem-statement mutations on SWE-bench Pro, systematically varying the information available to an agent while holding the underlying task fixed. We remove or isolate selected bug-report content and related task information, delete fault-localization cues, and test structural changes that flatten lists or remove section headers without changing the text itself. By measuring how each change affects agent solve rates, we find that both models depend on localization cues and expected behavior, and that structural changes alone can reduce solve rates, even without removing any content. The two models diverge in how they handle missing information: Qwen searches more widely and can exhaust its turn budget, while Gemma commits to a plausible interpretation early and patches on it.  \nOur findings indicate that a good bug report for an agent overlaps with, but is not identical to, a good report for a human: agents benefit most from concrete, executable, and well-localized information, whereas some qualities long emphasized for human readers, such as natural language steps to reproduce and readable descriptions, contribute little or even correlate with lower success.  \nI. INTRODUCTION  \nWriting a bug report is often the first step toward resolving a software defect. The report, written in natural language, lays out the problem and guides the investigation that follows. Reporters include information they believe will help identify the cause of the issue, localize the fault, and ultimately produce a fix. A good bug report helps the issue get resolved faster [1] . LLM-based automated program repair (APR) agents have moved into real-world engineering workflows, where they read a bug report and submit a patch that attempts to resolve the issue [2],[3] . Yet they still start from the same artifact humans use: the bug report. The agent works through tools that let it read files, run code, and search the repository, with each result  \nbecoming part of the context for its next decision. The report determines what information the agent has at the start and shapes the investigation that follows.  \nThis raises the question: What makes a good bug report for an AI repair agent? A line of empirical software engineering research has studied what makes a bug report useful for humans [4]–[6] . Prior work has examined which report characteristics developers find most valuable and how report quality relates to triage time, resolution time, and the overall effec","cbCaipcktfT6pXGA","https://ap.wps.com/l/cbCaipcktfT6pXGA","pdf",256413,2,1,12,"English","en",105,"# Introduction\n## Research motivation\n## Two complementary studies","[{\"question\":\"为什么传统的“好Bug报告”可能不一定适用于AI修复代理？\",\"answer\":\"论文指出，人类开发者与AI代理在信息需求上存在差异：AI代理无法像人类那样追问或协作补全缺失信息，因此报告缺失的部分往往需要代理自行恢复或直接缺失。\"},{\"question\":\"研究如何评估Bug报告特征与修复成功之间的关系？\",\"answer\":\"研究包含两部分：使用SWE-bench Verified中87个代理对433个问题的尝试结果，标注27类既有Bug报告特征，并用混合效应模型估计各特征对解决概率的影响。\"},{\"question\":\"控制消融实验的核心结论是什么？\",\"answer\":\"消融结果表明两种模型都强依赖定位线索与符合预期的行为信息；即使不移除文本内容，结构性变化也可能降低解决率。同时，Qwen倾向更广泛搜索并可能耗尽轮次预算，而Gemma会更早做出看似合理的解释并基于该解释进行修补。\"}]",1784186411,30,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"what-makes-a-good-bug-report-for-an-ai-agent","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/what-makes-a-good-bug-report-for-an-ai-agent/83269/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-23","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"为什么传统的“好Bug报告”可能不一定适用于AI修复代理？","Question",{"text":75,"@type":76},"论文指出，人类开发者与AI代理在信息需求上存在差异：AI代理无法像人类那样追问或协作补全缺失信息，因此报告缺失的部分往往需要代理自行恢复或直接缺失。","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"研究如何评估Bug报告特征与修复成功之间的关系？",{"text":80,"@type":76},"研究包含两部分：使用SWE-bench Verified中87个代理对433个问题的尝试结果，标注27类既有Bug报告特征，并用混合效应模型估计各特征对解决概率的影响。",{"name":82,"@type":73,"acceptedAnswer":83},"控制消融实验的核心结论是什么？",{"text":84,"@type":76},"消融结果表明两种模型都强依赖定位线索与符合预期的行为信息；即使不移除文本内容，结构性变化也可能降低解决率。同时，Qwen倾向更广泛搜索并可能耗尽轮次预算，而Gemma会更早做出看似合理的解释并基于该解释进行修补。","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":29,"slug":121},"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]