[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86124-en":3,"doc-seo-86124-105":30,"detail-sidebar-cat-0-en-105":83},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86124,962075114765,"Quinn","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","MJ Multi-turn LLM Jailbreaking via Decomposed Credit Assignment","Modern large language models are used in interactive multi-turn conversations, making multi-turn jailbreaks a realistic safety threat and a key target for automated red teaming. Learning effective multi-turn jailbreak attackers depends on accurate credit assignment, since different turns contribute unequally to the final harmful outcome. DC-GRPO assigns turn-level group-relative learning signals by combining immediate and future credit, preventing errors caused by coarse trajectory-level scoring. Experiments on multiple victim LLMs and benchmarks report ASR5@3 of 98.26% (dynamic) and 97.88% (static), surpassing SEMA and TROJail.","MJ: Multi-turn LLM Jailbreaking via Decomposed  \nCredit Assignment  \nJunyoung Park  \nPOSTECH GSAI  \n[pjy0422@postech.ac.kr](pjy0422@postech.ac.kr)  \nNamgyu Park  \nSamsung SDS [nam9yu.park@samsung.com](nam9yu.park@samsung.com)  \nSechan Lee  \nPOSTECH GSAI  \n[chan1031@postech.ac.kr](chan1031@postech.ac.kr)  \nYoon-Chan Jhi  \nSamsung SDS [yoonchan.jhi@samsung.com](yoonchan.jhi@samsung.com)  \nJihoon Cho  \nSamsung SDS [jihoon1.cho@samsung.com](jihoon1.cho@samsung.com)  \nSangdon Park  \nPOSTECH GSAI & CSE [sangdon@postech.ac.kr](sangdon@postech.ac.kr)  \narXiv :2607 . 1 1070v 1 [ cs .CL] 13 Jul 2026  \nAbstract  \nModern large language models (LLMs) operate in interactive multi-turn settings, making multi-turn jailbreaking a realistic threat model and an important setting for automated red teaming. A core challenge in learning multi-turn jailbreak attackers is credit assignment: different turns contribute differently to the final outcome, yet existing learning signals are often too coarse to identify their individual contributions. We propose decomposed credit GRPO (DC-GRPO), a unified turn-level credit assignment framework for Group Relative Policy Optimization in multiturn jailbreak learning. DC-GRPO assigns a separate group-relative learning signal to each turn by combining immediate and future credit, avoiding the credit misassignment induced by broadcasting a single trajectory-level score across the dialogue. We instantiate this framework with static and dynamic weighting rules that differ in how the two credit sources are balanced while sharing the same turnlevel structure. Across multiple victim LLMs and benchmarks, the dynamic-and static-weighted variants achieve average ASR5 @3 scores of 98.26% and 97.88%, respectively, substantially outperforming the state-of-the-art methods, including SEMA (86.58%) and TROJail (86.23%) . Their consistently strong performance indicates that the central empirical benefit comes from turn-level group-relative credit assignment rather than a particular weighting rule.  \n* Warning: This paper contains examples of harmful content.  \n1 Introduction  \nModern Large Language Models (LLMs) [1, 2, 3] are increasingly deployed in multi-turn conversational settings, where users and models exchange information over multiple turns rather than through a single isolated prompt. Safety alignment techniques such as RLHF [4, 5] have made these systems substantially more resistant to direct harmful requests, but vulnerabilities remain in multi-turn dialogue, where an attacker can strategically shape context, disguise intent, and adapt to the victim model’s responses across turns. Studying jailbreaks in this interactive setting is therefore important for realistic safety evaluation and for automated red teaming, where the goal is to identify multi-turn failure modes and ultimately improve model robustness against malicious use.  \nEffective automated red teaming requires attackers that are both scalable and adaptive, properties that neither hand-crafted prompts nor large closed models can fully provide. The former demands substantial human expertise with limited coverage [6, 7], while the latter incurs prohibitive computational cost and poor reproducibility [8, 9] . This creates a practical need for learned attacker policies  \nPreprint.  \nthat can be trained once and deployed repeatedly at scale, enabling even relatively small models to explore diverse and effective multi-turn attack trajectories.  \nTraining such policies, however, is non-trivial. A central challenge in learning such multi-turn jailbreak attackers is credit assignment. In multi-turn dialogue, not all attacker turns play the same role: some turns produce immediate progress toward a jailbreak, while others mainly prepare context that only becomes useful several turns later. Existing training-based methods [10, 11] have begun to learn attacker policies beyond hand-crafted prompting, but their learning signals remain too coarse for multi-turn dialogue. In ","cbCaivTnZrBEflfP","https://ap.wps.com/l/cbCaivTnZrBEflfP","pdf",872094,3,1,29,"English","en",105,"# Abstract\n# Warning\n# Introduction\n## Multi-turn jailbreak threat model\n## Automated red teaming requirements\n## Credit assignment challenge in multi-turn learning\n## Proposed approach: DC-GRPO","[{\"question\":\"How do the static and dynamic weighting variants differ?\",\"answer\":\"Static weighting uses a fixed coefficient to control the balance between future credit sources, while dynamic weighting derives coefficients from rollout-group statistics without introducing an extra mixing hyperparameter. Both share the same turn-level credit structure.\"}]",1784208675,73,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":78,"head_meta":80,"extra_data":82,"updated_unix":28},"mj-multi-turn-llm-jailbreaking-via-decomposed-credit-assignment","",{"@graph":36,"@context":77},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/mj-multi-turn-llm-jailbreaking-via-decomposed-credit-assignment/86124/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-26","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71],{"name":72,"@type":73,"acceptedAnswer":74},"How do the static and dynamic weighting variants differ?","Question",{"text":75,"@type":76},"Static weighting uses a fixed coefficient to control the balance between future credit sources, while dynamic weighting derives coefficients from rollout-group statistics without introducing an extra mixing hyperparameter. Both share the same turn-level credit structure.","Answer","https://schema.org",{"og:url":51,"og:type":79,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":81,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":84},[85,89,93,97,102,107,112,115,120,123,127],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":86,"show_sort_weight":87,"slug":88},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":90,"show_sort_weight":91,"slug":92},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Exam",70,"exam",{"id":98,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},5,"Comic",60,"comic",{"id":103,"doc_module":4,"doc_module_name":46,"category_name":104,"show_sort_weight":105,"slug":106},6,"Technology",50,"technology",{"id":108,"doc_module":4,"doc_module_name":46,"category_name":109,"show_sort_weight":110,"slug":111},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":113,"slug":114},30,"research-report",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},9,"Religion & Spirituality",20,"religion-spirituality",{"id":118,"doc_module":4,"doc_module_name":46,"category_name":121,"show_sort_weight":118,"slug":122},"World Cup","world-cup",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":124,"slug":126},10,"Lifestyle","lifestyle",{"id":128,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":98,"slug":130},19,"General","general"]