[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82849-en":3,"doc-seo-82849-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},82849,2336464648322,"Aria","https://ap-avatar.wpscdn.com/avatar/2200025388227c56fec?_k=1778556882303663488",8,"Research & Report","RUSTMIZAN: A Compilable, Contamination-Aware Benchmarking Framework for Rust Vulnerabilities","LLM agents are increasingly used for vulnerability analysis, yet current benchmarks often use small non-compilable snippets, emphasize only vulnerable-vs-nonvulnerable binary detection, and overlook dataset contamination from models being trained on released corpora. RUSTMIZAN provides a Rust vulnerability benchmarking framework with compilable code variants at crate, file, and function levels, annotations for binary detection, CWE classification, and function- and line-level localization. A semantics-preserving mutation framework enables contamination testing and robustness probing, showing competitive binary accuracy while line localization remains low and degrades under adversarial cues.","arXiv :2607 .04729v 1 [ cs .CR] 6 Jul 2026  \nRUSTMIZAN: A Compilable, Contamination-Aware Benchmarking Framework for Rust Vulnerabilities  \nTarek Elsayed1 ∗ Shiping Yang1 Eunsong Koh1 Sanika Goyal1 Vincent Huang 1 Paul Ngo 1 Nathan Young 1 Mohammad Omidvar Tehrani 1 Alvyn Kang 1 Arnell Kang 1 Zeyu Chen2 Angélica Moreira3  \nXuan Feng3 Angel X. Chang 1,4 Nick Sumner 1 Steven Y. Ko 1  \n1 Simon Fraser University 2Trinity University 3Microsoft Research 4Amii  \n[https://sfu-rsl.github.io/rust-mizan/](https://sfu-rsl.github.io/rust-mizan/)  \nAbstract  \nLLM agents are increasingly applied to vulnerability analysis, but existing benchmarks have not kept pace. They typically rely on small non-compilable snippets, focus on binary classification (vulnerable or not), and do not account for the risk that publicly-released datasets are part of model training corpora. We introduce RUSTMIZAN, a benchmarking framework for Rust vulnerability analysis that addresses these gaps. RUSTMIZAN contains compilable code variants at the crate, file, and function levels, with annotations for binary vulnerability detection, CWE classification, and function-and line-level localization. A paired mutation framework produces semantics-preserving code mutants for contamination testing and robustness probing. Across four frontier models in an agentic setup with commandline access, binary classification sits in the 56–65% range, but line localization F1 stays near 20%, and adversarial cues drop line F1 by about 27% .  \n1 Introduction  \nAI agents powered by large language models (LLMs) increasingly drive vulnerability analysis, giving researchers and practitioners capabilities they previously lacked. These agents analyze entire codebases as well as individual code snippets, compile and run programs, and effectively understand and use external tools. Beyond identifying vulnerabilities, they let users examine many related properties, such as vulnerability classes and affected functions or modules. Moreover, model vendors continuously train the underlying LLMs on extensive corpora of source code from hosting platforms such as GitHub.  \nExisting benchmarks, however, have not kept pace with these capabilities. Many recent benchmarks for AI agents target general software development [Jimenez et al., 2024, Merrill et al., 2026, Zanet al., 2026], which differs fundamentally from vulnerability analysis. On the other hand, existing vulnerability benchmarks often rely on small code snippets drawn from version-control diffs, stripping away the build environment necessary to compile and execute the code. They typically focus on binary detection (vulnerable or not), neglecting related properties such as vulnerability class and relevant contextual information. This setup cannot exercise the capabilities thorough analysis requires, e.g., using tools such as static analyzers and fuzzers, interacting with real build systems and codebases, and iterating on the analysis [Sheng et al., 2025, Ding et al., 2025] . They also suffer from data contamination: once a dataset is released, LLMs are trained on it, and the models then recall the vulnerabilities from memory instead of using their reasoning abilities [Cao et al., 2026, Ding et al., 2025, Wu et al., 2023] .  \n∗ Correspondence: [tareknaser360@gmail.com](tareknaser360@gmail.com)  \nPreprint.  \nTo address these limitations, we introduce RUSTMIZAN (mizan in Arabic means scale or balance), a benchmarking framework for Rust vulnerability analysis. We focus on Rust as security-critical infrastructure increasingly adopts it, including the Linux kernel, Android, and Windows, yet it remains underrepresented in vulnerability benchmarks [Cao et al., 2026] . Using Rust as a concrete example, RUSTMIZAN showcases the following three design principles that, to the best of our knowledge, no other vulnerability benchmarks provide in combination, including Rust benchmarks [Qin et al., 2020, Xu et al., 2021, Zheng et al., 2023, Androutsopoulos and Bianc","cbCainrEFPotYQIg","https://ap.wps.com/l/cbCainrEFPotYQIg","pdf",2100806,2,1,36,"English","en",105,"# Abstract\n# Introduction\n## Motivation: gaps in existing benchmarks\n## RUSTMIZAN design principles\n### Multi-task evaluation\n### Multi-level compilable variants\n### Contamination mitigation and robustness testing","[{\"question\":\"What problems does RUSTMIZAN aim to solve in existing vulnerability benchmarks?\",\"answer\":\"It targets three gaps: benchmarks often use non-compilable snippets, focus mainly on binary vulnerability detection, and do not account for contamination when models have been trained on released datasets.\"},{\"question\":\"Which evaluation tasks does RUSTMIZAN support?\",\"answer\":\"RUSTMIZAN supports four complementary tasks: binary vulnerability classification, CWE classification, function-level localization, and line-level localization, covering detection through precise diagnosis.\"},{\"question\":\"How does RUSTMIZAN test robustness against memorization and adversarial cues?\",\"answer\":\"It includes automated semantics-preserving code mutants for contamination testing and additional mutants that introduce adversarial cues such as misleading comments, allowing probing of model reliance on reasoning versus recall.\"}]",1784183404,91,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"rustmizan-a-compilable-contamination-aware-benchmarking-framework-for-rust-vulnerabilities","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/rustmizan-a-compilable-contamination-aware-benchmarking-framework-for-rust-vulnerabilities/82849/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problems does RUSTMIZAN aim to solve in existing vulnerability benchmarks?","Question",{"text":75,"@type":76},"It targets three gaps: benchmarks often use non-compilable snippets, focus mainly on binary vulnerability detection, and do not account for contamination when models have been trained on released datasets.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Which evaluation tasks does RUSTMIZAN support?",{"text":80,"@type":76},"RUSTMIZAN supports four complementary tasks: binary vulnerability classification, CWE classification, function-level localization, and line-level localization, covering detection through precise diagnosis.",{"name":82,"@type":73,"acceptedAnswer":83},"How does RUSTMIZAN test robustness against memorization and adversarial cues?",{"text":84,"@type":76},"It includes automated semantics-preserving code mutants for contamination testing and additional mutants that introduce adversarial cues such as misleading comments, allowing probing of model reliance on reasoning versus recall.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]