[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84252-en":3,"doc-seo-84252-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},84252,13056703019662,"Evangeline","https://ap-avatar.wpscdn.com/avatar/be000253a8e92610077?_k=1778726343310543188",8,"Research & Report","Rethinking Code Performance Benchmarks for LLMs","Function-level performance benchmarks are widely used to test whether large language models (LLMs) generate efficient programs, yet benchmark results often show little or no execution-time advantage over canonical solutions. This study revisits four benchmarks—EffiBench, Enamel, EvalPerf, and Mercury—by executing 1,538 tasks 30 times and applying statistical tests to compare runtimes. Only 6.11% of benchmark-reported performant implementations are significantly faster; manual review indicates many improvements remain hidden by original test suites.","arXiv :2607 .076 19v 1 [ cs . SE] 8 Jul 2026  \nNoname manuscript No.  \n(will be inserted by the editor)  \nRethinking Code Performance Benchmarks for LLMs  \nNhat Minh Le · Yisen Xu · Zhijie  \nWang · Tse-Hsun (Peter) Chen  \nReceived: date / Accepted: date  \nAbstract Many function-level performance benchmarks have been proposed to evaluate whether large language models (LLMs) can generate efficient programs. However, results on these benchmarks often show that LLM-generated implementations have little or no execution-time difference from canonical solutions. This observation motivates us to revisit these benchmarks and examine whether they are suitable for performance evaluation. In this paper, werevisit four popular benchmarks: EffiBench, Enamel, EvalPerf, and Mercury. We evaluate 1,538 tasks under more rigorous setting by running each task 30 times and assessing the runtime differences between the canonical solutionsand benchmark-provided performant implementations with statistical testing. With the benchmark-provided test suites, only 6 . 11% of the performant implementations are significantly faster than the canonical solutions. In a manual analysis of 308 non-significant tasks, 99 performant implementations contain no meaningful performance change, while 209 contain potential performance improvements that are not exposed by the original tests.  \nNhat Minh Le  \nSoftware Performance, Analysis, and Reliability (SPEAR) Lab, Concordia University, Montreal, Quebec, Canada  \nE-mail: [nhatminh.le@mail.concordia.ca](nhatminh.le@mail.concordia.ca)  \nYisen Xu  \nSoftware Performance, Analysis, and Reliability (SPEAR) Lab, Concordia University, Montreal, Quebec, Canada  \nE-mail: [yisen.xu@mail.concordia.ca](yisen.xu@mail.concordia.ca)  \nZhijie Wang  \nConcordia University, Montreal, Quebec, Canada E-mail: [zhijie.wang@concordia.ca](zhijie.wang@concordia.ca)  \n[Tse-Hsun](Tse-Hsun) ([Peter](Peter)) [Chen](Chen)  \nSoftware Performance, Analysis, and Reliability (SPEAR) Lab, Concordia University, Montreal, Quebec, Canada  \nE-mail: [peterc@encs.concordia.ca](peterc@encs.concordia.ca)  \nThese results suggest that the main limitation is not only the evaluation method, but also the limited sufficiency of the benchmark-provided performance tests. To address this limitation, we propose an LLM-based multi-agent framework to generate performance-oriented tests that expose runtime differences more effectively than the original tests. The framework uses three separate agents to generate, diagnose, and repair deterministic tests that preserve functional correctness while better exposing performance differences. Across 1,345 benchmark tasks for which the original tests found no significant performance difference, tests generated by our framework with DeepSeek-v3.1 and GPT-4o reveal statistically significant improvements in 24.01% and 25.43% of the tasks, respectively, outperforming the SOTA LLM-based performance test generation method. Finally, we discuss the implications for future performance benchmark for LLM-generated code. In addition to using repeated execution and statistical testing to improve rigor, future work should consider selecting problems with meaningful opportunities for performance optimization rather than relying on overly simple tasks, constructing sufficiently challenging testcases that make runtime differences between implementations observable, and extending this line of evaluation from isolated function-level tasks to class-level or repository-level settings where performance bottlenecks may arise from interactions among multiple components.  \nKeywords Performance Benchmark, Large Language Models, Code Generation  \n1 Introduction  \nSoftware performance has a long research history in the software engineering community (Woodside et al., 2007) . Decades of work have studied how to measure, diagnose (Jin et al., 2012; Baltes et al., 2015), and improve runtime behavior through performance testing (Vokolos and Weyuker, 1998; Weyuker and","cbCaiah9ApaX43Gp","https://ap.wps.com/l/cbCaiah9ApaX43Gp","pdf",636623,4,1,38,"English","en",105,"# Introduction\n## Motivation and background\n## Prior benchmarks and evaluation settings\n## Function-level efficiency benchmarks","[{\"question\":\"What problem do the authors identify with existing LLM code performance benchmarks?\",\"answer\":\"They find that benchmark outcomes frequently show minimal or no execution-time difference versus canonical solutions, suggesting the benchmarks may not be suitable for performance evaluation.\"},{\"question\":\"How do the authors evaluate the four benchmarks in this paper?\",\"answer\":\"They run 1,538 tasks under more rigorous conditions, executing each task 30 times and using statistical testing to compare runtimes between canonical solutions and benchmark-provided performant implementations.\"},{\"question\":\"What improvement does the proposed multi-agent framework provide?\",\"answer\":\"It generates more performance-oriented deterministic tests that better expose runtime differences, yielding statistically significant performance improvements in 24.01% (DeepSeek-v3.1) and 25.43% (GPT-4o) of tasks where original tests found no significant difference.\"}]",1784194377,96,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"rethinking-code-performance-benchmarks-for-llms","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/rethinking-code-performance-benchmarks-for-llms/84252/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem do the authors identify with existing LLM code performance benchmarks?","Question",{"text":75,"@type":76},"They find that benchmark outcomes frequently show minimal or no execution-time difference versus canonical solutions, suggesting the benchmarks may not be suitable for performance evaluation.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How do the authors evaluate the four benchmarks in this paper?",{"text":80,"@type":76},"They run 1,538 tasks under more rigorous conditions, executing each task 30 times and using statistical testing to compare runtimes between canonical solutions and benchmark-provided performant implementations.",{"name":82,"@type":73,"acceptedAnswer":83},"What improvement does the proposed multi-agent framework provide?",{"text":84,"@type":76},"It generates more performance-oriented deterministic tests that better expose runtime differences, yielding statistically significant performance improvements in 24.01% (DeepSeek-v3.1) and 25.43% (GPT-4o) of tasks where original tests found no significant difference.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]