[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86155-en":3,"doc-seo-86155-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86155,962075114101,"Seraphina","https://ap-avatar.wpscdn.com/avatar/e000253a75eb197efd?x-image-process=image/resize,m_fixed,w_180,h_180&k=1780044092746381165",8,"Research & Report","AMT-X: A Phase-Structured Multi-Turn Red-Teaming Framework with Checklist-Gated Dual-Metric Evaluation for LLM Safety","LLM safety evaluation often relies on single-turn attack datasets and single-judge scoring, which underestimates risk from adaptive adversaries that operate through extended dialogues. AMT-X (Adaptive Multi-Turn Exploitation) introduces a phase-structured multi-turn red-teaming framework as an explicit, reproducible state machine with semantic-driven transitions. It replaces single-judge evaluation with a multi-role jury and phase-conditioned checklists, yielding dual metrics that distinguish partially actionable outputs from fully operational, real, and actionable harm.","arXiv :2607 . 1 1 15 1v 1 [ cs .CR] 13 Jul 2026  \nAMT-X: A Phase-Structured Multi-Turn Red-Teaming Framework with Checklist-Gated Dual-Metric Evaluation for LLM Safety  \nYI TING SHEN, Vulcan Research, AIFT, Singapore KENTAROH TOYODA, Vulcan Research, AIFT, Singapore ALEX LEUNG, Vulcan Research, AIFT, Singapore  \nSafety evaluation of large language models (LLMs) relies largely on single-turn attack datasets and single-judge scoring, underestimating risk from adaptive multi-turn adversaries and reporting a single success rate that does not separate partially actionable outputs from those carrying complete operational detail. We propose AMT-X (Adaptive Multi-Turn Exploitation), a phase-structured multi-turn red-teaming framework. Unlike prior multi-turn attacks that rely on ad hoc escalation or free-form per-goal plans, AMT-X casts the attack as an explicit, reproducible multi-phase state machine driven by semantic signals from the victim, and replaces single-judge scoring with a multi-role jury whose phase-conditioned checklists gate success on actionable harm. Across six frontier victim models (queried under their default safety alignment, without added moderation layers) and seven Moderation sub-categories, AMT-X attains overall attack success rates of 97.6–100% under a lenient score threshold, but 66.7–78.6% under a stricter gate requiring complete, real, and operational detail: a gap of up to 33 percentage points between partially and fully actionable harm.  \nCCS Concepts: • Security and privacy → Software and application security; • Computing methodologies → Natural language processing.  \nAdditional Key Words and Phrases: LLM safety, red-teaming, multi-turn jailbreak, adversarial evaluation  \n1 Introduction  \nThe rapid deployment of large language models in production environments (spanning medical advisory systems, customer service agents, autonomous coding assistants, and general-purpose chatbots) has created an urgent need for rigorous adversarial evaluation. Safety fine-tuning through reinforcement learning from human feedback (RLHF) [25] and constitutional AI [4] has significantly reduced single-turn harmful outputs, but the adversarial landscape has evolved in response. Attackers no longer rely on a single carefully crafted prompt; instead, they engage in extended, adaptive conversations that build context, exploit logical inconsistencies, and progressively extract policy-violating outputs. As attackers become more sophisticated, the central question becomes whether our evaluation methodology keeps pace.  \nMainstream safety benchmarks have not kept pace with this shift. HarmBench [22], AdvBench [41], and JailbreakBench [6] measure robustness against static single-turn datasets, providing strong comparability across models but systematically understating risk from adaptive multi-turn adversaries. Iterative attacks such as PAIR [7] and TAP [23], multi-turn methods such as Crescendo [33] and Chain of Attack (CoA) [38], and more recent agentic attackers such as GOAT [27], ActorAttack [32], and X-Teaming [31] have demonstrated that adaptive and conversational strategies can substantially bypass safety-aligned models; human multi-turn red-teamers similarly defeat defenses that report single-digit ASRs under automated single-turn attacks [15] . However, these methods still share three structural limitations that we revisit in detail in Section 3: (i) attacks are improvised turn by turn, so trajectories are neither reproducible nor structurally attributable: two runs of the “same” attack can diverge, and it remains unclear which components actually drive success; (ii) success is typically reported asa single LLM score, which does not separate nominal successes that carry complete operational detail from those that stop short of it; and (iii) the same single LLM judge that grades success also exhibits documented positional, verbosity, and self-preference biases, further undermining reported numbers.  \nAuthors’ Contact Informat","cbCaieAYJyPGspcz","https://ap.wps.com/l/cbCaieAYJyPGspcz","pdf",1971465,10,1,35,"English","en",105,"# Introduction\n## Motivation for multi-turn evaluation\n## Limitations of existing benchmarks\n## Proposed AMT-X contributions","[{\"question\":\"Why do existing LLM safety benchmarks understate multi-turn attack risk?\",\"answer\":\"They mainly use static single-turn datasets and single-judge scoring, which do not reflect adaptive conversations that accumulate context and progressively extract policy-violating outputs.\"},{\"question\":\"How does AMT-X make multi-turn red-teaming more reproducible and attributable?\",\"answer\":\"AMT-X casts attacks as an explicit, reproducible multi-phase state machine with semantic-signal-driven phase transitions and a structured technique library.\"},{\"question\":\"What are the two evaluation metrics in AMT-X, and why are they important?\",\"answer\":\"AMT-X computes a lenient success rate and a stricter full success rate, separating partially actionable outputs from fully operational detail, so the gap between them becomes measurable.\"}]",1784208961,88,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"amt-x-a-phase-structured-multi-turn-red-teaming-framework-with-checklist-gated-dual-metric-evaluation-for-llm-safety","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/amt-x-a-phase-structured-multi-turn-red-teaming-framework-with-checklist-gated-dual-metric-evaluation-for-llm-safety/86155/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"Why do existing LLM safety benchmarks understate multi-turn attack risk?","Question",{"text":76,"@type":77},"They mainly use static single-turn datasets and single-judge scoring, which do not reflect adaptive conversations that accumulate context and progressively extract policy-violating outputs.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does AMT-X make multi-turn red-teaming more reproducible and attributable?",{"text":81,"@type":77},"AMT-X casts attacks as an explicit, reproducible multi-phase state machine with semantic-signal-driven phase transitions and a structured technique library.",{"name":83,"@type":74,"acceptedAnswer":84},"What are the two evaluation metrics in AMT-X, and why are they important?",{"text":85,"@type":77},"AMT-X computes a lenient success rate and a stricter full success rate, separating partially actionable outputs from fully operational detail, so the gap between them becomes measurable.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":20,"slug":134},"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":107,"slug":138},19,"General","general"]