[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83555-en":3,"doc-seo-83555-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83555,34359740700684,"Finn","https://ap-avatar.wpscdn.com/avatar/1f400023980c374ae676?_k=1777273430885731487",8,"Research & Report","SWE-Doctor: Guiding Software Engineering Agents with Runtime Diagnosis from Multi-Faceted Bug Reproduction Tests","LLM-based software engineering agents generate patches from issue reports and code repositories, yet bug reproduction tests (BRTs) are not clearly understood as guidance for patch generation. A preliminary study shows that using advanced BRT generators directly is ineffective: fail-to-fail BRTs can mislead agents, while even fail-to-pass BRTs yield limited or negative gains due to partial coverage. SWE-Doctor instead uses multi-faceted BRT executions to build runtime-grounded diagnosis records, improving patch guidance and reducing partial patches. Evaluations on SWE-bench Verified and SWE-bench Pro across five LLM backends show consistent superiority and higher resolution rates.","SWE-Doctor: Guiding Software Engineering Agents with Runtime Diagnosis from Multi-Faceted Bug  \nReproduction Tests  \nYaoqi Guo∗ , Yang Liu∗ , Jie M. Zhang†, Yun Ma‡, Yiling Lou§ , Zhenpeng Chen¶  \n∗ Nanyang Technology University, †King’s College London, ‡Peking University, § University of Illinois Urbana-Champaign,  \n¶ Tsinghua University  \n[yaoqi001@e.ntu.edu.sg](yaoqi001@e.ntu.edu.sg), [yangliu@ntu.edu.sg](yangliu@ntu.edu.sg), [jie.zhang@kcl.ac.uk](jie.zhang@kcl.ac.uk), [mayun@pku.edu.cn](mayun@pku.edu.cn), [yilingl@illinois.edu](yilingl@illinois.edu),  \n[zpchen@tsinghua.edu.cn](zpchen@tsinghua.edu.cn)  \narXiv :2607 .00990v 1 [ cs . SE] 1 Jul 2026  \nAbstract—Large language model (LLM)-based software engineering agents are increasingly developed to resolve software issues by generating patches from issue reports and code repositories. Bug reproduction tests (BRTs) are an important building block for such agents and have been shown useful for patch validation. However, it remains unclear whether BRTs can also help the more central stage of patch generation. We first conduct a preliminary study and find that directly using advanced BRT generators to guide patch generation is not beneficial: fail-to-fail BRTs can mislead agents, while even fail-topass BRTs bring limited or negative gains. Our analysis reveals two reasons: fail-to-pass BRTs may cover only one manifestation of the reported issue, leading to partial patches, whereas failto-fail BRTs are unreliable as direct patch-generation targets. Motivated by these insights, we propose SWE-Doctor, a software issue resolution agent that guides patch generation with runtime diagnoses derived from multi-faceted BRT executions. SWEDoctor first generates multi-faceted BRTs for different behavioral requirements stated in the issue, then executes and debugs these BRTs to construct runtime-grounded diagnosis records, and finally uses the diagnoses together with localization information inferred during BRT generation to guide patch generation and reduce partial patches. We evaluate SWE-Doctor on Python bug-fixing issues from the widely adopted SWE-bench Verified and SWE-bench Pro across five LLM backends. SWE-Doctor consistently outperforms existing agents across all 10 LLM– benchmark combinations, achieving average resolution rates of 75.7% on SWE-bench Verified and 59.4% on SWE-bench Pro. In particular, on the more challenging SWE-bench Pro, SWE-Doctor improves the average resolution rate by 8.0–8.9 percentage points over the baseline agents. Further analyses show that both multi-faceted BRT generation and runtime-grounded diagnosis contribute substantially to SWE-Doctor’s effectiveness.  \nIndex Terms—Software Issue Resolution, Agent, Bug Reproduction Test, Runtime Diagnosis  \nI. INTRODUCTION  \nLarge language models (LLMs) have accelerated the development of software engineering agents for resolving realworld software issues [1]–[3] . Given a reported issue and its corresponding buggy repository, such agents can navigate the codebase, inspect relevant files, edit source code, run tests, and  \nCorresponding author: Zhenpeng Chen.  \nsubmit patches in an end-to-end manner, with LLMs serving as their reasoning and decision-making backends.  \nBug reproduction test (BRT) generation is the task of generating tests that reproduce the bugs described in issue reports [4] . It is an important building block for software issue resolution because a BRT turns a textual bug report into executable feedback [1], [5] . In real-world issue resolution tasks, however, such tests are usually not provided [2]: an agent must infer the bug-triggering behavior from the issue report and the repository itself. As a result, effective BRT generation can provide agents with concrete executions that expose the reported bug and support downstream resolution.  \nRecent studies [4], [6], [7] have proposed advanced BRT generators and shown that generated BRTs can improve issue resolution when used for patch validation, i.","cbCaikvtqulpIf4J","https://ap.wps.com/l/cbCaikvtqulpIf4J","pdf",1258859,5,1,12,"English","en",105,"# Introduction\n## Bug reproduction tests in agent-based patch generation\n## Limitations of direct BRT guidance\n## Proposed approach: runtime-grounded diagnoses\n# Evaluation","[{\"question\":\"Why are bug reproduction tests important for LLM-based patch generation agents?\",\"answer\":\"Bug reproduction tests turn a textual bug report into executable feedback by exposing bug-triggering behavior through concrete executions, which can support downstream resolution.\"},{\"question\":\"What problems arise when directly using advanced BRT generators to guide patch generation?\",\"answer\":\"Fail-to-pass BRTs may cover only one manifestation of the reported issue, leading to partial patches, while fail-to-fail BRTs can mislead agents as direct patch-generation targets.\"},{\"question\":\"How does SWE-Doctor guide patch generation differently from prior BRT usage?\",\"answer\":\"SWE-Doctor generates multi-faceted BRTs, executes and debugs them to produce runtime-grounded diagnosis records, and then uses these diagnoses together with localization information to guide patch generation and reduce partial patches.\"}]",1784188798,30,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"swe-doctor-guiding-software-engineering-agents-with-runtime-diagnosis-from-multi-faceted-bug-reproduction-tests","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/swe-doctor-guiding-software-engineering-agents-with-runtime-diagnosis-from-multi-faceted-bug-reproduction-tests/83555/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-26","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"Why are bug reproduction tests important for LLM-based patch generation agents?","Question",{"text":76,"@type":77},"Bug reproduction tests turn a textual bug report into executable feedback by exposing bug-triggering behavior through concrete executions, which can support downstream resolution.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"What problems arise when directly using advanced BRT generators to guide patch generation?",{"text":81,"@type":77},"Fail-to-pass BRTs may cover only one manifestation of the reported issue, leading to partial patches, while fail-to-fail BRTs can mislead agents as direct patch-generation targets.",{"name":83,"@type":74,"acceptedAnswer":84},"How does SWE-Doctor guide patch generation differently from prior BRT usage?",{"text":85,"@type":77},"SWE-Doctor generates multi-faceted BRTs, executes and debugs them to produce runtime-grounded diagnosis records, and then uses these diagnoses together with localization information to guide patch generation and reduce partial patches.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,110,115,120,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":29,"slug":121},"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":20,"slug":137},19,"General","general"]