[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-81888-en":3,"doc-seo-81888-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":11,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},81888,2336464648322,"Aria","https://ap-avatar.wpscdn.com/avatar/2200025388227c56fec?_k=1778556882303663488",8,"Research & Report","Refused in Chat, Written in Code Workflow-Level Jailbreak Construction in IDE Coding Agents","Large language models increasingly power IDE-integrated coding agents that decompose tasks, edit files, execute code, and refine outputs over many turns, but safety is often assessed like a chatbot using isolated prompts. This work introduces workflow-level jailbreak construction, where a harmful objective is assembled across ordinary software-development workflow stages. Using GitHub Copilot in VS Code and four closed backends, the study shows near-complete refusal in direct-chat baselines yet 816/816 unsafe completions under the full multi-turn workflow, confirmed by expert evaluators.","Refused in Chat, Written in Code: Workflow-Level Jailbreak Construction in IDE Coding Agents.  \nAbhishek Kumar∗ , Carsten Maple†  \nThe Alan Turing Institute, London, United Kingdom  \n∗ [akumar@turing.ac.uk](akumar@turing.ac.uk)  \n†[cmaple@turing.ac.uk](cmaple@turing.ac.uk), [cm@warwick.ac.uk](cm@warwick.ac.uk)  \narXiv :2607 .03968v2 [ cs . SE] 9 Jul 2026  \nAbstract—Large language models are increasingly deployed as IDE-integrated coding agents that decompose tasks, generate and edit files, run code, and refine outputs over many turns. Yet their safety is still often evaluated as if they were chatbots: one harmful prompt, one response, judged in isolation. We introduce workflowlevel jailbreak construction, a failure mode in which a harmful objective is assembled across ordinary stages of a softwaredevelopment workflow rather than generated through a single direct prompt. Using GitHub Copilot in Visual Studio Code, we study four closed-weight backends: Claude Sonnet 4.6, Claude Haiku 4.5, Gemini 3.1 Pro, and Gemini 3.5 Flash. Across 204 prompts from Hammurabi’s Code, HarmBench, and AdvBench, the models show near-complete refusal under direct chat, CSVread, and single-step code-fix baselines, with only 8/816 successful responses in each baseline condition. Under the full workflow, however, the same prompts and backends produce 816/816 unsafe teaching-shot completions, all independently confirmed by two expert evaluators under a strict rubric. These results show that conversational refusal benchmarks can substantially overstate the safety of deployed coding agents and motivate defenses that reason about safety across multi-turn IDE workflows and their generated artifacts, not only individual chat turns.  \nIndex Terms—Coding agents, IDE security, Jailbreak attacks, LLM safety, Agentic AI.  \nI. INTRODUCTION  \nLarge language models (LLMs) have rapidly reshaped software engineering, moving from passive code-completion engines to active participants in the software-development loop [1]–[5] . Through IDE-integrated coding agents, these systems interpret developer goals, generate and edit files, execute commands, read execution results, debug failures, andrefine their output across many turns [6]–[8] . This shift is not merely a capability upgrade; it changes the shape of the safety problem [9], [10] . A safety failure is no longer necessarily a single harmful prompt answered by a single harmful response. It can instead emerge gradually, distributed across a multi-turn development workflow in which every individual interaction looks like an ordinary programming task.  \nExisting safety evaluation has only partially adapted to the shift from chatbots to agentic systems. Much of the jailbreak literature still evaluates models through prompt-level or conversation-level interactions. Prompt-level attacks, including optimization-based [11], [12], in-context [13], [14], and codereformulation attacks [15], can bypass safety behavior in frontier, closed-weight models. Multi-turn conversational attacks have also been demonstrated [16] . These studies show that refusal behavior can fail in general LLM settings, but they  \ndo not fully capture the software-development workflows in which coding agents operate. In parallel, software-engineering work has begun to study harmful behavior in coding contexts. Code Red introduces Hammurabi’s Code, a benchmark of software-engineering-specific harmful prompts spanning malware, copyright misuse, and other dangerous tasks [17] . Another recent study shows that coding agents can fail during ordinary development by breaking constraints, performing destructive actions, optimizing the wrong metric, or falsely claiming that a task has been completed [18] . What remainsunexamined is whether the IDE coding-agent workflow itself, including task decomposition, scripting, execution, and metricdriven refinement, can be turned into the attack surface. This gap is especially important because prior software-engineering safe","cbCaibON38EXat3x","https://ap.wps.com/l/cbCaibON38EXat3x","pdf",5758285,1,11,"English","en",105,"# Introduction\n## Workflow-level jailbreak construction\n## Experimental setup and evaluation results","[{\"question\":\"What is “workflow-level jailbreak construction” in IDE coding agents?\",\"answer\":\"It is a failure mode where a harmful objective is assembled across multiple ordinary stages of a software-development workflow, rather than being produced by a single direct prompt.\"},{\"question\":\"How do the models behave under baseline chat-style evaluations versus the full workflow?\",\"answer\":\"Under direct chat/CSV/single-step code-fix baselines, the models largely refuse or give safety-aligned outputs; under the full multi-turn workflow, the same objectives lead to unsafe teaching-shot completions.\"},{\"question\":\"What evidence supports the workflow-level vulnerability?\",\"answer\":\"Across 204 prompts and four closed-weight backends, baseline conditions yield only 8/816 successful unsafe responses, while the full workflow yields 816/816 unsafe completions, independently confirmed by two expert evaluators under a strict rubric.\"}]","Refused in Chat, Written in Code Workflow-Level Jailbreak Construction in IDE Coding Agents | PDF",1784176882,28,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"refused-in-chat-written-in-code-workflow-level-jailbreak-construction-in-ide-coding-agents","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/refused-in-chat-written-in-code-workflow-level-jailbreak-construction-in-ide-coding-agents/81888/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-31","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":11},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What is “workflow-level jailbreak construction” in IDE coding agents?","Question",{"text":76,"@type":77},"It is a failure mode where a harmful objective is assembled across multiple ordinary stages of a software-development workflow, rather than being produced by a single direct prompt.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How do the models behave under baseline chat-style evaluations versus the full workflow?",{"text":81,"@type":77},"Under direct chat/CSV/single-step code-fix baselines, the models largely refuse or give safety-aligned outputs; under the full multi-turn workflow, the same objectives lead to unsafe teaching-shot completions.",{"name":83,"@type":74,"acceptedAnswer":84},"What evidence supports the workflow-level vulnerability?",{"text":85,"@type":77},"Across 204 prompts and four closed-weight backends, baseline conditions yield only 8/816 successful unsafe responses, while the full workflow yields 816/816 unsafe completions, independently confirmed by two expert evaluators under a strict rubric.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":46,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":46,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]