[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84997-en":3,"doc-seo-84997-105":29,"detail-sidebar-cat-0-en-105":90},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":11,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},84997,13056703019404,"Miles","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","Beyond Attack Success Rate: Action Graded Severity Scale for Tool Using AI Agents","Agentic red-teaming benchmarks often summarize prompt-injection outcomes with a single attack-success bit, failing to reflect what defenders most need: how harmful the agent’s executed actions were. This work proposes an action-graded harm rubric on a seven-level ordinal scale (L0–L6) driven by tool-call trajectories, scoring reversibility, scope crossing, privilege expansion, and escalation chains. Severity is computed via a deterministic oracle and validated against a three-judge LLM panel, exposing cases where binary ASR is misleading.","Beyond Attack-Success Rate: Action-Graded Severity Scale for Tool-Using AI Agents  \nHarry Owiredu-Ashley  \nIndependent Researcher  \nNew Jersey, USA  \nowireduashlh1@montclair.edu  \narXiv :2607 .07474v 1 [ cs .CR] 8 Jul 2026  \nAbstract—Agentic red-teaming benchmarks report whether an injected agent was compromised as a single bit: the attack succeeded, or it did not. We argue that this binary attack-success rate discards the information a defender most needs, namely how harmful the resulting action was. We introduce an action-graded harm rubric that scores an agent’s tool-call trajectory on a sevenlevel ordinal scale (L0 to L6) according to whether the executed action was reversible, whether it crossed scope to reach another party, and whether it expanded privilege. We compute the scale two ways: a deterministic oracle that reads the trajectory and the attacker’s stated goal, and a panel of three frontier languagemodel judges that read a tag-free account of the same trajectory. Across four victim models and two defenses on the AgentDojo workspace suite, severity grading exposes three cases the binary metric hides, including a defense that reports a zero attacksuccess rate while still permitting a cross-scope leak through an unfiltered tool. The judge panel reproduces the oracle with high ordinal agreement (Krippendorff’s α = 0 .91) but shares systematic blind spots that we characterize, most notably a failure to recognize escalation chains. Unlike prior work that provides harm taxonomies, harmful-task completion tests, execution-level safety benchmarks, or severity-aware simulation, our contribution is a reusable, trace-grounded severity instrument applied to the actual actions recorded in existing red-team logs. All code, prompts, and per-episode logs are released.  \nIndex Terms—large language models, agentic AI, prompt injection, red teaming, AI safety, severity scoring, LLM-as-judge, tool-using agents  \nI. INTRODUCTION Language-model agents no longer only produce text.  \nEquipped with tools, they read mailboxes, send email, edit files, and grant access to shared resources. This shift moves the safety question from what the model says to what the agent does. A prompt injection buried in a retrieved document is no longer only a matter of eliciting a disallowed sentence; it is a matter of causing an irreversible, externally visible action in the environment the agent operates on [1], [2] .  \nAgentic red-teaming benchmarks capture this threat with a binary outcome: for each episode the injected task either succeeded or it did not, and the headline metric is the resulting attack-success rate (ASR) [1], [3] . This is the same reduction that we and others have criticized in the single-turn conversational setting, where a binary jailbreak label discards the trajectory signal that a sustained interaction provides [4],[5], [6] . In the agentic setting the discarded signal is different, and arguably more consequential: it is the severity of the action the agent actually took.  \nConsider two episodes that a benchmark records identically as “attack succeeded.” In the first, the agent saved an attackersupplied note to a file it alone can read. In the second, the agent shared a sensitive document with an external party, granting them standing access. Both are one bit. A defender deciding whether to deploy an agent needs to distinguish them, and the binary metric cannot. Conversely, a defense can drive the reported ASR to zero while still permitting a cross-scope action through a tool whose effect the benchmark’s success check happened not to watch. The bit is not merely coarse; it can be affirmatively misleading.  \nWe propose to grade the outcome of an agentic episode by the worst action the agent actually executed, on a sevenlevel ordinal scale ( L0 to L6) that we call an action-graded harm rubric. The scale is defined by three effect axes that a defender cares about independently of the attacker’s intent: whether the action was reversible","cbCainieqKWxLrMC","https://ap.wps.com/l/cbCainieqKWxLrMC","pdf",382055,3,1,"English","en",105,"# Introduction\n## Motivation: Limits of binary attack-success rate\n## Action-graded harm rubric (L0–L6)\n## Two severity computation methods: oracle and LLM judges\n## Evaluation on AgentDojo and key contributions","[{\"question\":\"Why is a binary attack-success rate insufficient for agentic red-teaming?\",\"answer\":\"Because it collapses the trajectory into one bit, losing information about the actual harm level of the action the agent executed. Two episodes can share the same ASR while one is reversible and limited and another grants access to another party.\"},{\"question\":\"How does the action-graded harm rubric determine severity levels?\",\"answer\":\"It assigns an ordinal level from L0 to L6 based on three effect axes: reversibility, scope crossing (reaching another party or shared state), and privilege expansion. An additional level captures escalation chains across steps.\"},{\"question\":\"How are severity labels computed and validated in this work?\",\"answer\":\"A deterministic oracle derives severity from raw tool-call trajectories using per-tool effect metadata and an attribution rule tied to the attacker’s stated goal. A panel of three frontier language-model judges independently scores a tag-free natural-language account of the same trajectory, achieving high ordinal agreement with the oracle (Krippendorff’s α = 0.91).\"}]",1784200124,20,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":85,"head_meta":87,"extra_data":89,"updated_unix":27},"beyond-attack-success-rate-action-graded-severity-scale-for-tool-using-ai-agents","",{"@graph":35,"@context":84},[36,52,67],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,49],{"item":40,"name":41,"@type":42,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":20},"https://docshare.wps.com/document/research-report/",{"item":50,"name":13,"@type":42,"position":51},"https://docshare.wps.com/document/beyond-attack-success-rate-action-graded-severity-scale-for-tool-using-ai-agents/84997/",4,{"url":50,"name":13,"@type":53,"author":54,"headline":13,"publisher":56,"fileFormat":59,"inLanguage":23,"description":14,"dateModified":60,"datePublished":61,"encodingFormat":59,"isAccessibleForFree":62,"interactionStatistic":63},"DigitalDocument",{"name":9,"@type":55},"Person",{"url":40,"name":57,"@type":58},"DocShare","Organization","application/pdf","2026-07-23","2026-07-16",true,{"@type":64,"interactionType":65,"userInteractionCount":20},"InteractionCounter",{"@type":66},"ViewAction",{"@type":68,"mainEntity":69},"FAQPage",[70,76,80],{"name":71,"@type":72,"acceptedAnswer":73},"Why is a binary attack-success rate insufficient for agentic red-teaming?","Question",{"text":74,"@type":75},"Because it collapses the trajectory into one bit, losing information about the actual harm level of the action the agent executed. Two episodes can share the same ASR while one is reversible and limited and another grants access to another party.","Answer",{"name":77,"@type":72,"acceptedAnswer":78},"How does the action-graded harm rubric determine severity levels?",{"text":79,"@type":75},"It assigns an ordinal level from L0 to L6 based on three effect axes: reversibility, scope crossing (reaching another party or shared state), and privilege expansion. An additional level captures escalation chains across steps.",{"name":81,"@type":72,"acceptedAnswer":82},"How are severity labels computed and validated in this work?",{"text":83,"@type":75},"A deterministic oracle derives severity from raw tool-call trajectories using per-tool effect metadata and an attribution rule tied to the attacker’s stated goal. A panel of three frontier language-model judges independently scores a tag-free natural-language account of the same trajectory, achieving high ordinal agreement with the oracle (Krippendorff’s α = 0.91).","https://schema.org",{"og:url":50,"og:type":86,"og:title":13,"og:site_name":57,"og:description":14},"article",{"robots":88,"canonical":50},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":91},[92,96,100,104,109,114,119,122,126,129,133],{"id":21,"doc_module":4,"doc_module_name":45,"category_name":93,"show_sort_weight":94,"slug":95},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":97,"show_sort_weight":98,"slug":99},"Literature",80,"literature",{"id":51,"doc_module":4,"doc_module_name":45,"category_name":101,"show_sort_weight":102,"slug":103},"Exam",70,"exam",{"id":105,"doc_module":4,"doc_module_name":45,"category_name":106,"show_sort_weight":107,"slug":108},5,"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":45,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":28,"slug":125},9,"Religion & Spirituality","religion-spirituality",{"id":28,"doc_module":4,"doc_module_name":45,"category_name":127,"show_sort_weight":28,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":45,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":45,"category_name":135,"show_sort_weight":105,"slug":136},19,"General","general"]