[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83293-en":3,"doc-seo-83293-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},83293,1374391974564,"Clementine","https://ap-avatar.wpscdn.com/avatar/14000253aa45c000a9e?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779874745381141002",8,"Research & Report","ScopeJudge: Cost-Aware Pre-Execution Gating for Offensive Security Agents","As LLM agents perform offensive security tasks, a single out-of-scope tool call can violate engagement boundaries, disrupt production, or invalidate bug-bounty results. Unlike fixed safety policies, the correct boundary is declared by the user’s request and must be inferred from intent. ScopeJudge studies pre-execution gating using a cheap LLM judge that accepts or rejects proposed calls before execution. The work introduces a labeled benchmark of 4,897 tool calls and evaluates judge models and context strategies to map cost–accuracy trade-offs.","ScopeJudge: Cost-Aware Pre-Execution Gating for Offensive  \nSecurity Agents  \nShane Caldwell∗ dreadnode, USA  \nMax Harley† dreadnode, USA  \nAds Dawson‡ dreadnode, USA  \nMichael Kouremetis§ dreadnode, USA  \nVincent Abruzzo¶ dreadnode, USA  \nWill Pearce ‖ dreadnode, USA  \narXiv :2607 .07774v 1 [ cs .CR] 8 Jul 2026  \nAbstract  \nAs LLM agents take on offensive security work, a single out-of-scope tool call can breach a client’s engagement boundary, disrupt production, or void a bug-bounty finding. Unlike a fixed safety policy, the boundary that matters is declared in the user’s request and must be inferred from intent. That challenge is sharpened by the adversarial nature of offensive security: the same tool call is in or out of scope depending not on the action itself but on the target it touches and the context in which it runs, which no fixed policy can enumerate in advance. We study pre-execution gating: a cheap, trusted LLM judge inspects each call proposed by a strong, swappable agent, and accepts or rejects it before it runs. We introduce ScopeJudge, a benchmark of 4,897 tool calls (7 .7% scope violations) from agent trajectories on tasks engineered to tempt agents out of scope and labeled at the call level by professional penetration testers, with substantial inter-grader agreement (Fleiss κ = 0 .64) that sets an expert agreement reference point of F1 = 0 .78. We evaluate eight judge models under five transcript strategies, varying how much context the judge sees, from the static policy alone to the full raw transcript, and chart the resulting cost–accuracy Pareto frontier. We find that a static policy is structurally insufficient for scope enforcement:  \nblind to the user’s request, judge recall collapses to near zero, confirming that scope lives in the request and that request-conditioned monitoring is necessary. We find the strongest judges are open-weight: GLM-5.2 reaches F1 = 0 .66 , the highest of any judge we test, beating the best proprietary judge (0 .60) at roughly one-third the per-call cost. Because a missed violation costs more than a spurious rejection, we report precision, recall, and F1 separately and recommend two operating points: a cost-sensitive configuration anda recall-first one for high-stakes deployments. We release the ScopeJudge dataset to support real-time monitoring and scalable oversight of autonomous security agents.  \n1 Introduction  \nLLMs now act as agents: given an objective and a set of tools, they plan and issue tool calls autonomously until a long-horizon goal is met [1 , 2] . With this autonomy has come new risk. Practitioners have already reported software-engineering agents that delete production databases and other cloud infrastructure irreversibly [3] . It is easy to imagine how this would  \n∗ dreadnode, Principal Research Engineer. Email: [shane@dreadnode.io | GitHub: @SJCaldwell](shane@dreadnode.io | GitHub: @SJCaldwell)[ ](shane@dreadnode.io | GitHub: @SJCaldwell)†dreadnode, Principal [Security Researcher. Email: max@dreadnode.io | GitHub: @t94j0](Security Researcher. Email: max@dreadnode.io | GitHub: @t94j0)  \n‡dreadnode, [Staff AI Security Researcher. Email: ads@dreadnode.io | GitHub: @GangGreenTemperTatum](Staff AI Security Researcher. Email: ads@dreadnode.io | GitHub: @GangGreenTemperTatum)[ ](Staff AI Security Researcher. Email: ads@dreadnode.io | GitHub: @GangGreenTemperTatum)§dreadnode, [Principal AI Research Engineer. Email: michael@dreadnode.io | GitHub: @mkultraWasHere](Principal AI Research Engineer. Email: michael@dreadnode.io | GitHub: @mkultraWasHere)[ ](Principal AI Research Engineer. Email: michael@dreadnode.io | GitHub: @mkultraWasHere)¶ dreadnode, Principal Research Engineer. Email: [vincent@dreadnode.io | GitHub: @vabruzzo](vincent@dreadnode.io | GitHub: @vabruzzo)  \n‖ dreadnode, Co-Founder. Email: [will@dreadnode.io | GitHub: @moohax](will@dreadnode.io | GitHub: @moohax)  \nappend to next  \nreject (1)  \nReject / escalate  \nFigure 1: Pre-execution gating. At eac","cbCaiqs0SR4rh1je","https://ap.wps.com/l/cbCaiqs0SR4rh1je","pdf",769709,1,22,"English","en",105,"# Abstract\n# Introduction","[{\"question\":\"Why can out-of-scope tool calls be especially risky for offensive security agents?\",\"answer\":\"A tool call that is outside the engagement boundary can breach client terms, disrupt production systems, or undermine a bug-bounty finding. The risk comes from scope being defined by the user’s request rather than a fixed action rule.\"},{\"question\":\"What problem does ScopeJudge address and how does it work?\",\"answer\":\"ScopeJudge introduces pre-execution gating: a low-cost, trusted LLM judge inspects each proposed tool call from a stronger agent and accepts or rejects it before execution. This uses intent-conditioned context to decide whether the call is in scope.\"},{\"question\":\"How are judge models evaluated and what findings are reported?\",\"answer\":\"Eight judge models are evaluated under five transcript strategies that vary the amount of context provided to the judge, producing a cost–accuracy Pareto frontier. A static policy is shown to be structurally insufficient, with the best open-weight judge reaching the highest reported F1 while improving cost compared to the best proprietary judge.\"}]",1784186531,55,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"scopejudge-cost-aware-pre-execution-gating-for-offensive-security-agents","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/scopejudge-cost-aware-pre-execution-gating-for-offensive-security-agents/83293/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why can out-of-scope tool calls be especially risky for offensive security agents?","Question",{"text":75,"@type":76},"A tool call that is outside the engagement boundary can breach client terms, disrupt production systems, or undermine a bug-bounty finding. The risk comes from scope being defined by the user’s request rather than a fixed action rule.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What problem does ScopeJudge address and how does it work?",{"text":80,"@type":76},"ScopeJudge introduces pre-execution gating: a low-cost, trusted LLM judge inspects each proposed tool call from a stronger agent and accepts or rejects it before execution. This uses intent-conditioned context to decide whether the call is in scope.",{"name":82,"@type":73,"acceptedAnswer":83},"How are judge models evaluated and what findings are reported?",{"text":84,"@type":76},"Eight judge models are evaluated under five transcript strategies that vary the amount of context provided to the judge, producing a cost–accuracy Pareto frontier. A static policy is shown to be structurally insufficient, with the best open-weight judge reaching the highest reported F1 while improving cost compared to the best proprietary judge.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]