[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84590-en":3,"doc-seo-84590-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},84590,16904993612988,"Olivia Brown","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","LLVM-Bench Benchmarking and Advancing Large Language Models for LLVM Compiler Issue Resolution","LLVM is a widely used compiler infrastructure whose scale and complexity make issue resolution labor-intensive and challenging. While large language models (LLMs) have shown strong performance in software issue resolution, their effectiveness on complex system-level LLVM compiler issues is insufficiently studied. This work introduces LLVMBench, a large-scale benchmark with 423 validated LLVM tasks, and LLVM-Gym, an automated evaluation platform. Experiments across multiple LLMs, retrieval settings, and agents reveal dominant failure modes and motivate a lightweight ensemble, LLVM-Ens.","LLVM-Bench: Benchmarking and Advancing Large Language Models for LLVM Compiler Issue  \nResolution  \nZhao Tian  \nSchool of Computer Software, Tianjin University, Tianjin, China [tianzhao@tju.edu.cn](tianzhao@tju.edu.cn)  \nYingquan Zhao  \nSchool of Cybersecurity, Tianjin University, Tianjin, China [zhaoyingquan@tju.edu.cn](zhaoyingquan@tju.edu.cn)  \nChenyao Suo  \nSchool of Computer Software, Tianjin University, Tianjin, China [chenyaosuo@tju.edu.cn](chenyaosuo@tju.edu.cn)  \narXiv :2607 .00700v 1 [ cs . SE] 1 Jul 2026  \nMeng Wang  \nSchool of Computer Science, University of Bristol, Bristol, UK [meng.wang@bristol.ac.uk](meng.wang@bristol.ac.uk)  \nJunjie Chen*  \nSchool of Computer Software, Tianjin University, Tianjin, China [junjiechen@tju.edu.cn](junjiechen@tju.edu.cn)  \nAbstract—LLVM is a widely used compiler infrastructure whose scale and complexity make issue resolution labor-intensive and challenging. Although large language models (LLMs) have recently achieved remarkable success in issue resolution, their effectiveness on complex system-level LLVM compiler remains largely unexplored. To address this gap, we introduce LLVMBench, the first large-scale benchmark for LLVM issue resolution, containing 423 real-world, validated tasks collected from the LLVM project. We further develop LLVM-Gym, a scalable evaluation platform that automates issue reproduction, patch application, compiler building, and test execution. Using LLVMBench and LLVM-Gym, we conduct a comprehensive study of four representative LLMs, six retrieval configurations, and three agents. Our results show that current LLM-based issue resolution techniques remain limited on LLVM-Bench, with patch invalidity and build failures as the dominant failure modes. We further reveal a strong complementarity among different LLMs and agents, motivating LLVM-Ens, a lightweight ensemble approach that expands the patch space through integrating the patches generated by diverse techniques, filters incorrect and redundant candidates, and identifies the most promising solution. Our results show that LLVM-Ens achieves a resolution rate of up to 21.99%, further improving LLVM issue resolution.  \nI. INTRODUCTION  \nLLVM compiler infrastructure is one of the most influential and widely adopted software systems in modern computing [1]–[3] . It provides a collection of modular and reusable compiler and toolchain technologies that support a broad range of programming languages [4]–[7] . Owing to its flexibility, extensibility, and high performance, LLVM has become a foundational component in both industrial and academic compiler and program analysis projects [8]–[13] . As LLVM continues to evolve, however, developers must continuously address newly discovered bugs and feature requests. To date, LLVM community has accumulated nearly 100,000 issues on GitHub,  \n*Junjie Chen is the corresponding author.  \nof which approximately 30,000 remain unresolved [14] . Resolving these issues often requires understanding a largescale and highly complex codebase, implementing non-trivial code changes, and preserving existing functionality while addressing the target issue [15]–[17] . Consequently, LLVM issue resolution remains labor-intensive and costly, motivating the development of automated issue resolution techniques to improve both developer productivity and software quality.  \nRecent advances in large language models (LLMs) have demonstrated remarkable capabilities in software issue resolution [18]–[20] . Building upon the planning, reasoning, and tool-use abilities of LLMs, coding agents can autonomously navigate software repositories, execute commands, and generate patches for real-world issues. These capabilities have led to substantial progress on issue resolution benchmarks, most notably SWE-bench [15], a widely used benchmark that evaluates the ability of LLMs to resolve application-level software issues. Specifically, LLMs are provided only with an issue description and the corresponding codeb","cbCaipUD5I12mz42","https://ap.wps.com/l/cbCaipUD5I12mz42","pdf",1609278,3,1,12,"English","en",105,"# Abstract\n# Introduction\n## Motivation: challenges in LLVM issue resolution\n## Limits of existing benchmarks\n## Contributions: LLVMBench and LLVM-Gym","[{\"question\":\"What problem does LLVM-Bench address?\",\"answer\":\"LLVM-Bench targets the gap in evaluating LLM-based techniques on complex, system-level LLVM compiler issue resolution, which existing benchmarks do not adequately cover.\"},{\"question\":\"What are the main components introduced in the paper?\",\"answer\":\"The paper introduces LLVMBench, a benchmark with 423 validated LLVM tasks, and LLVM-Gym, a platform that automates issue reproduction, patch application, compiler building, and test execution.\"},{\"question\":\"What failure modes are reported as dominant on LLVM-Bench?\",\"answer\":\"Experiments indicate patch invalidity and build failures are the dominant failure modes for current LLM-based issue resolution approaches.\"}]",1784196962,30,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"llvm-bench-benchmarking-and-advancing-large-language-models-for-llvm-compiler-issue-resolution","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/llvm-bench-benchmarking-and-advancing-large-language-models-for-llvm-compiler-issue-resolution/84590/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-21","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does LLVM-Bench address?","Question",{"text":75,"@type":76},"LLVM-Bench targets the gap in evaluating LLM-based techniques on complex, system-level LLVM compiler issue resolution, which existing benchmarks do not adequately cover.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What are the main components introduced in the paper?",{"text":80,"@type":76},"The paper introduces LLVMBench, a benchmark with 423 validated LLVM tasks, and LLVM-Gym, a platform that automates issue reproduction, patch application, compiler building, and test execution.",{"name":82,"@type":73,"acceptedAnswer":83},"What failure modes are reported as dominant on LLVM-Bench?",{"text":84,"@type":76},"Experiments indicate patch invalidity and build failures are the dominant failure modes for current LLM-based issue resolution approaches.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":29,"slug":121},"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]