[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84937-en":3,"doc-seo-84937-105":30,"detail-sidebar-cat-0-en-105":82},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},84937,687197207639,"Asher","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","Reliable and Developer-Aligned Evaluation of Agents for Software Engineering","Large language models are closing the software development cycle, shifting from assistive tools to autonomous contributors within collaborative engineering workflows. Existing evaluation approaches remain fragmented and often rely on hypothetical syntactic setups, producing unreliable measurements and distorted estimates of true capability. This research proposes an end-to-end evaluation methodology for LLM-based agents grounded in real-world practice, emphasizing contamination awareness, in-the-wild agent behavior assessment, and trajectory-aware benchmarks that reflect coding context, human alignment, and model failure modes.","Reliable and Developer-Aligned Evaluation of Agents  \nfor Software Engineering  \nRazvan Mihai Popescu  \n[r.m.popescu@tudelft.nl](r.m.popescu@tudelft.nl)  \nDelft University of Technology  \nDelft, The Netherlands  \narXiv :2607 .067 13v 1 [ cs . SE] 7 Jul 2026  \nAbstract  \nLarge language models are rapidly moving towards closing the development cycle, transitioning from simple assistive companions to autonomous contributors deeply embedded into collaborative development environments. Despite their accelerated adoption, existing evaluation techniques are limited due to their fragmented nature and distorted projection of true model capabilities, often obtained from hypothetical syntactic scenarios. This research aims to bridge this gap by providing a comprehensive evaluation methodology for LLM-powered agents that is grounded in real-world software development practice. Our evaluation approach focuses on contamination-awareness, in-the-wild agentic behavior assessment, and trajectory-aware benchmarks and metrics capturing realistic coding contexts, human-aligned behavior, and model failure modes.  \nCCS Concepts  \n• Software and its engineering → Software creation and management.  \nKeywords  \nLarge Language Models (LLMs), Autonomous AI Agents, Software Engineering, Evaluation, Benchmarking, Data Contamination  \nACM Reference Format:  \nRazvan Mihai Popescu. 2026. Reliable and Developer-Aligned Evaluation of Agents for Software Engineering. In 34th ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering (FSE Companion ’26), July 05–09, 2026, Montreal, QC, Canada. ACM, New York, NY, USA, 2 pages. [https://doi.org/10.1145/3803437.3804877](https://doi.org/10.1145/3803437.3804877)  \n1 Introduction  \nGiven their strong reasoning and abstraction capabilities, Large Language Models (LLMs) have been widely employed in the context of Software Engineering (SE), which has become a prominent application domain. These models have shown a remarkable capacity to understand the structured nature of source code and achieve strong performance in various code-related tasks, leading to an increasing reliance on empirical evaluations to assess their abilities and limitations.  \nHowever, current evaluation practices are often unreliable, difficult to reproduce, and poorly aligned with real-world developer  \nThis work is licensed under a Creative Commons Attribution-NonCommercialNoDerivatives 4 .0 International License.  \nFSE Companion ’26, Montreal, QC, Canada  \n© 2026 Copyright held by the owner/author(s) .  \nACM ISBN 979-8-4007-2636-1/2026/07  \n[https://doi.org/10.1145/3803437.3804877](https://doi.org/10.1145/3803437.3804877)  \nneeds. Many works strongly rely on convenience-based metrics borrowed from NLP, such as BLEU, METEOR, or ROUGE, which fail to capture functional validity and developer utility [2, 7] . At the sametime, many of these metrics are inconsistently used across different coding tasks with minimal to no adaptation, while their coding variants, such as CodeBLEU, or embedding-based versions, such as CodeBERTScore, have shown to perform on par with general translation metrics [2, 7] . Furthermore, these metric-level limitations are amplified by flaws in widely used benchmarks. Various benchmarks suffer from issues such as saturation, data contamination, incorrect ground truths, and unrealistic context scenarios, which not only distort the reported performance, but also hinder the comparability and reproducibility of the studies [3] . Lastly, although LLMs are increasingly being used as evaluation judges, their assessments are prone to biases and hallucinations, often diverging from human judgements and further compromising the validity of evaluation protocols [8] .  \nNevertheless, these evaluation practices are still the norm in the community, despite often creating an illusion of model competence and resulting in misleading conclusions and potentially harmful downstream effects. The recent","cbCaiaC8aG3L6Ev3","https://ap.wps.com/l/cbCaiaC8aG3L6Ev3","pdf",380283,3,1,2,"English","en",105,"# Introduction\n# Related Work","[{\"question\":\"What limitations are associated with common code evaluation metrics and benchmarks?\",\"answer\":\"Metrics borrowed from NLP (e.g., BLEU, METEOR, ROUGE) fail to represent functional validity and developer utility, and their coding variants often show limited adaptation. Benchmarks may become saturated or contaminated, harming comparability and reproducibility.\"}]",1784199553,5,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":77,"head_meta":79,"extra_data":81,"updated_unix":28},"reliable-and-developer-aligned-evaluation-of-agents-for-software-engineering","",{"@graph":36,"@context":76},[37,52,67],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,49],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":22},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":50,"name":13,"@type":43,"position":51},"https://docshare.wps.com/document/reliable-and-developer-aligned-evaluation-of-agents-for-software-engineering/84937/",4,{"url":50,"name":13,"@type":53,"author":54,"headline":13,"publisher":56,"fileFormat":59,"inLanguage":24,"description":14,"dateModified":60,"datePublished":61,"encodingFormat":59,"isAccessibleForFree":62,"interactionStatistic":63},"DigitalDocument",{"name":9,"@type":55},"Person",{"url":41,"name":57,"@type":58},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":64,"interactionType":65,"userInteractionCount":20},"InteractionCounter",{"@type":66},"ViewAction",{"@type":68,"mainEntity":69},"FAQPage",[70],{"name":71,"@type":72,"acceptedAnswer":73},"What limitations are associated with common code evaluation metrics and benchmarks?","Question",{"text":74,"@type":75},"Metrics borrowed from NLP (e.g., BLEU, METEOR, ROUGE) fail to represent functional validity and developer utility, and their coding variants often show limited adaptation. Benchmarks may become saturated or contaminated, harming comparability and reproducibility.","Answer","https://schema.org",{"og:url":50,"og:type":78,"og:title":13,"og:site_name":57,"og:description":14},"article",{"robots":80,"canonical":50},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":83},[84,88,92,96,100,105,110,113,118,121,125],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":85,"show_sort_weight":86,"slug":87},"Story & Novel",90,"story-novel",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":89,"show_sort_weight":90,"slug":91},"Literature",80,"literature",{"id":51,"doc_module":4,"doc_module_name":46,"category_name":93,"show_sort_weight":94,"slug":95},"Exam",70,"exam",{"id":29,"doc_module":4,"doc_module_name":46,"category_name":97,"show_sort_weight":98,"slug":99},"Comic",60,"comic",{"id":101,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},6,"Technology",50,"technology",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":111,"slug":112},30,"research-report",{"id":114,"doc_module":4,"doc_module_name":46,"category_name":115,"show_sort_weight":116,"slug":117},9,"Religion & Spirituality",20,"religion-spirituality",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":119,"show_sort_weight":116,"slug":120},"World Cup","world-cup",{"id":122,"doc_module":4,"doc_module_name":46,"category_name":123,"show_sort_weight":122,"slug":124},10,"Lifestyle","lifestyle",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":127,"show_sort_weight":29,"slug":128},19,"General","general"]