[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84489-en":3,"doc-seo-84489-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},84489,549758146520,"Patrick","https://ap-avatar.wpscdn.com/avatar/80002397d8c0411e94?_k=1775819394049821470",8,"Research & Report","CTFusion CTF-Based Benchmark for LLM Agent Evaluation via MCP","Recent advances in large language models enable agentic systems for complex, multi-step tasks, and cybersecurity has become a major use case. Capture The Flag (CTF) benchmarks are widely adopted, but existing CTF benchmarks reuse past challenges, creating data contamination risks and opportunities for cheating, especially when agents use web-search-based retrieval. CTFUSION proposes a streaming evaluation framework on LIVE CTFS, implemented as a Model Context Protocol (MCP) server on CTFD to preserve per-agent independence and limit leakage. Experiments show traditional benchmarks may be unreliable, while CTFUSION is robust.","CTFusion : A CTF Based Benchmark for Evaluating LLM Agents via MCP  \nDongjun Lee 1 Ga-eun Bae 1 Insu Yun 1  \narXiv :2605 . 1 1504v2 [ cs .LG] 11 Jul 2026  \nAbstract  \nRecent advances in Large Language Models (LLMs) have enabled agentic systems for complex, multi-step tasks; cybersecurity is emerging as a prominent application. To evaluate such agents, researchers widely adopt Capture The Flag (CTF) benchmarks. However, current CTF benchmarks reuse existing challenges, which exposes them to data contamination and potential cheating. Notably, we confirmed these issues in practice by integrating web search tools into an existing agent. To address these limitations, we present CTFUSION, a streaming evaluation framework built on LIVE CTFS. To achieve this, CTFUSION preserves per-agent independence under a single team account and reduces competition impact by forwarding only the first correct flag per challenge. Moreover, we implement CTFUSION as a Model Context Protocol (MCP) server on the widely used CTFD platform, which offers broad applicability to diverse CTF events and agent types. Through experiments with three LLMs, two agents, and five LIVE CTFS, we demonstrate that existing CTF benchmarks can be unreliable in assessing LLM-based agents, while CTFUSION can serve as a robust solution for evaluating cybersecurity agents. We release CTFUSION as opensource to foster future research in this area.  \n1. Introduction  \nRecent advances in LLM agents have substantially improved their capacity for complex, multi-step tasks. These advances have found significant applications in cybersecurity, particularly in software vulnerability discovery and automated exploitation. Notable examples include ENIGMA (Abramovich et al., 2025), D-CIPHER (Udeshi  \net al., 2025), and CRAKEN (Shao et al., 2025), which em- 1 School of Electrical Engineering, KAIST, Daejeon, Republic of Korea. Correspondence to: Insu Yun \u003C[insuyun@kaist.ac.kr](insuyun@kaist.ac.kr)>.  \nPublished at the Second Workshop on Agents in the Wild: Safety, Security, and Beyond (AIWILD) at ICML 2026 . Copyright 2026 by the author(s) .  \nploy LLM agents to automate cybersecurity tasks. Recently, XBOW (Waisman, 2025) has demonstrated its real-world impact by achieving the top rank on U.S. leaderboard of HackerOne, a well-known bug bounty platform.  \nTo evaluate these agents, CTF benchmarks have become the de-facto standard. CTF competitions present practical security challenges where participants need to uncover hidden flags. For instance, the NYU CTF BENCH (Shao et al., 2024b) was the first benchmark to use CTF problems for evaluating cybersecurity agents. It comprises adataset of CTF problems collected from competitions held between 2017 and 2023 . Moreover, CYBENCH (Zhang et al., 2025) carefully selected 40 professional-level CTF tasks from four distinct CTF competitions, chosen to be recent, meaningful, and spanning a wide range of difficulties. XBOW-BENCHMARK (Waisman, 2024) extends this CTF style and designs a benchmark to evaluate web-security tasks, reflecting real-world vulnerabilities. CTFKNOW (Ji et al., 2025) constructed 3,992 technical questions to assess LLMs’ knowledge in cybersecurity through CTF challenges. These studies suggest that while LLMs have a wealth of security knowledge, they struggle to effectively apply this knowledge to solve CTF problems.  \nDespite their usefulness, existing CTF benchmarks are inherently limited in fairly evaluating LLM-based agents for cybersecurity due to two main issues: data contamination and potential cheating. In particular, data contamination can arise when benchmark tasks overlap with training data, enabling models to memorize prior solutions. As LLM models have been continuously released, it is highly likely that old benchmarks were included in training corpora, making them vulnerable to data contamination. More seriously, agents often integrate Retrieval Augmented Generation (RAG) using web search to supplement their knowledge. Th","cbCailAsROdGv4RT","https://ap.wps.com/l/cbCailAsROdGv4RT","pdf",1068242,2,1,22,"English","en",105,"# Introduction\n## Limitations of existing CTF benchmarks\n## Proposed CTFUSION framework\n## Experimental setup and evaluation results","[{\"question\":\"Why can existing CTF benchmarks be unreliable for evaluating LLM agents?\",\"answer\":\"They can suffer from data contamination when challenges overlap with training data, and from potential cheating when agents use retrieval such as web search to obtain public write-ups or flags.\"},{\"question\":\"How does CTFUSION mitigate data contamination and cheating?\",\"answer\":\"CTFusion evaluates on LIVE CTFS with unreleased challenges and forwards only the first correct flag per challenge, while also preserving per-agent progress views under a single competition account.\"},{\"question\":\"How is CTFUSION implemented for use across CTF events?\",\"answer\":\"CTFusion is implemented as a Model Context Protocol (MCP) server on the widely used CTFD platform, enabling seamless integration of different agent types across diverse LIVE CTFS competitions.\"}]",1784196000,55,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"ctfusion-ctf-based-benchmark-for-llm-agent-evaluation-via-mcp","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/ctfusion-ctf-based-benchmark-for-llm-agent-evaluation-via-mcp/84489/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-21","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why can existing CTF benchmarks be unreliable for evaluating LLM agents?","Question",{"text":75,"@type":76},"They can suffer from data contamination when challenges overlap with training data, and from potential cheating when agents use retrieval such as web search to obtain public write-ups or flags.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does CTFUSION mitigate data contamination and cheating?",{"text":80,"@type":76},"CTFusion evaluates on LIVE CTFS with unreleased challenges and forwards only the first correct flag per challenge, while also preserving per-agent progress views under a single competition account.",{"name":82,"@type":73,"acceptedAnswer":83},"How is CTFUSION implemented for use across CTF events?",{"text":84,"@type":76},"CTFusion is implemented as a Model Context Protocol (MCP) server on the widely used CTFD platform, enabling seamless integration of different agent types across diverse LIVE CTFS competitions.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]