[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85575-en":3,"doc-seo-85575-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85575,1649267921044,"Ava Thompson","https://us-avatar.wpscdn.com/avatar/1800007509477c92dfb?_k=1782875107921204101",8,"Research & Report","HarDBench: Draft-Based Co-Authoring Jailbreak Attacks for Safe Human–LLM Collaborative Writing","Large language models increasingly act as co-authors in collaborative writing, where users start from rough drafts and ask the model to complete, revise, and refine text. This process can be abused: malicious users provide incomplete harmful drafts and jailbreak the model into producing dangerous completions. HarDBench is introduced as a systematic benchmark to evaluate robustness across high-risk domains, with realistic prompt structures and domain cues. A safety-utility balanced preference-optimization alignment reduces harmful outputs while preserving helpful co-authoring performance, establishing a new evaluation paradigm.","HarDBench: A Benchmark for Draft-Based Co-Authoring Jailbreak Attacks for Safe Human–LLM Collaborative Writing  \nEuntae Kim1 , Soomin Han2 , Buru Chang1 *  \n1 Korea University, 2 Sogang University  \n{untae0122,[buru_chang}@korea.ac.kr](buru_chang}@korea.ac.kr), [soominsion@u.sogang.ac.kr](soominsion@u.sogang.ac.kr)  \nWarning. This paper includes references to hazardous procedures, such as cyberattacks and explosives, solely to analyze and mitigate LLM vulnerabilities for research purposes.  \narXiv :2604 . 19274v 3 [ cs .CL] 13 Jul 2026  \nAbstract  \nLarge language models (LLMs) are increasingly used as co-authors in collaborative writing, where users begin with rough drafts and rely on LLMs to complete, revise, and refine their content. However, this capability poses a serious safety risk: malicious users could jailbreak the models—filling incomplete drafts with dangerous content—to force them into generating harmful outputs. In this paper, we identify the vulnerability of current LLMs to such draft-based co-authoring jailbreak attacks and introduce HarDBench, a systematic benchmark designed to evaluate the robustness of LLMs against this emerging threat. HarDBench spans a range of high-risk domains—including Explosives, Drugs, Weapons, and Cyberattacks—and features prompts with realistic structure and domain-specific cues to assess the model susceptibility to harmful completions. To mitigate this risk, we introduce a safety-utility balanced alignment approach based on preference optimization, training models to refuse harmful completions while remaining helpful on benign drafts. Experimental results show that existing LLMs are highly vulnerable in co-authoring contexts and our alignment method significantly reduces harmful outputs without degrading performance on co-authoring capabilities. This presents a new paradigm for evaluating and aligning LLMs in human-LLM collaborative writing settings.  \nOur new benchmark and dataset are available on our project page at [https://github.com/](https://github.com/)[ ](https://github.com/)untae0122/HarDBench  \n1 Introduction  \nLarge language models (LLMs) have demonstrated the ability to generate responses grounded in knowledge acquired from large-scale text corpora. As a result, many users now incorporate LLMs as coauthors in their writing processes, drawing upon  \n* Corresponding author.  \nthe models’ knowledge to complete and refine their writing (Lee et al., 2022 ; Noy and Zhang, 2023) .  \nIn particular, users often begin with a rough draft and employ LLMs to fill in missing knowledge, address argumentative gaps, and polish the text, thereby maximizing the model’s utility. Recent research on human preference optimization has improved LLMs’ co-authoring capabilities by utilizing preference data that capture aspects such as helpfulness, clarity, and writing quality (Ouyang et al., 2022 ; Ethayarajh et al., 2024) .  \nHowever, such draft-based co-authoring processes entail potential misuse. As shown in Figure 1, a malicious user can input an incomplete yet harmful draft (e.g., a partially written drug synthesis procedure) and prompt the model to polish it. Even with safety mechanisms in place, the LLM generates more harmful outputs from its internal knowledge, including detailed and executable instructions that could cause real-world harm. This exposes a significant risk: the model’s co-authoring capabilities can be exploited to surface harmful knowledge that would otherwise be restricted by system-level safeguards. Despite such risks, this issue remains largely unexplored in current research.  \nIn this paper, we propose HarDBench (Harmful Draft Benchmark), a benchmark grounded in the co-authoring process to systematically evaluate the vulnerabilities of LLMs. We begin by manually collecting representative domain-specific keywords (e.g., PETN, fentanyl, M16, Whonix) spanning four high-risk domains: Explosives, Drugs, Weapons, and Cyberattacks. Using these keywords, we prompt the drafter m","cbCaijUmx7kv1OTJ","https://ap.wps.com/l/cbCaijUmx7kv1OTJ","pdf",5223336,5,1,47,"English","en",105,"# Introduction\n## HarDBench Overview\n## Threat Model: Draft-Based Co-Authoring Jailbreaks\n## Dataset Construction and Benchmark Design\n## Safety–Utility Balanced Alignment\n## Experimental Findings","[{\"question\":\"What vulnerability does HarDBench focus on in human–LLM collaborative writing?\",\"answer\":\"HarDBench targets draft-based co-authoring jailbreaks, where a malicious user supplies an incomplete harmful draft and prompts the model to polish it, leading to harmful and detailed completions from internal knowledge.\"},{\"question\":\"Which high-risk domains are covered by HarDBench?\",\"answer\":\"HarDBench spans four high-risk domains: Explosives, Drugs, Weapons, and Cyberattacks, using representative domain-specific keywords and prompts that simulate realistic user behavior.\"},{\"question\":\"How does the proposed alignment method reduce harmful outputs without harming co-authoring ability?\",\"answer\":\"The approach uses preference optimization to balance safety and utility: harmful completions are treated as rejected while benign cooperative completions are preferred, training the model to refuse harmful drafts yet remain helpful on benign ones.\"}]",1784204697,118,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"hardbench-draft-based-co-authoring-jailbreak-attacks-for-safe-humanllm-collaborative-writing","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/hardbench-draft-based-co-authoring-jailbreak-attacks-for-safe-humanllm-collaborative-writing/85575/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What vulnerability does HarDBench focus on in human–LLM collaborative writing?","Question",{"text":76,"@type":77},"HarDBench targets draft-based co-authoring jailbreaks, where a malicious user supplies an incomplete harmful draft and prompts the model to polish it, leading to harmful and detailed completions from internal knowledge.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"Which high-risk domains are covered by HarDBench?",{"text":81,"@type":77},"HarDBench spans four high-risk domains: Explosives, Drugs, Weapons, and Cyberattacks, using representative domain-specific keywords and prompts that simulate realistic user behavior.",{"name":83,"@type":74,"acceptedAnswer":84},"How does the proposed alignment method reduce harmful outputs without harming co-authoring ability?",{"text":85,"@type":77},"The approach uses preference optimization to balance safety and utility: harmful completions are treated as rejected while benign cooperative completions are preferred, training the model to refuse harmful drafts yet remain helpful on benign ones.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":20,"slug":138},19,"General","general"]