[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83495-en":3,"doc-seo-83495-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83495,687197100911,"Himbo","https://ap-avatar.wpscdn.com/avatar/a000239b6f1da00475?x-image-process=image/resize,m_fixed,w_180,h_180&k=1782698725881665579",8,"Research & Report","Beyond the Prompt Jailbreaking Function Calling LLMs via Simulated Moderation Traces","Jailbreak attacks pose a serious risk to safe deployment of large language models, especially in stateful function-calling environments. The work identifies a structural weakness beyond prompt-level studies: trusted schemas, structured arguments, untrusted tool outputs, and accumulated dialog state are interleaved in one shared context, blurring control logic and data. SMT, a black-box simulated moderation-trace framework, builds a multi-turn trajectory that exploits validation feedback to progressively weaken safety constraints and induce harmful outputs. Extensive experiments on multiple commercial LLMs show consistently highest attack success and HarmScore with few queries, proving prompt sanitization is insufficient.","Beyond the Prompt: Jailbreaking Function-Calling LLMs via Simulated  \nModeration Traces  \nDisclaimer. This paper contains examples of harmful language. Reader discretion is recommended.  \nJunlong Liu 1,+ , Haobo Wang 1,+ , Weiqi Luo 1,* , Xiaojun Jia2  \n1Sun Yat-sen University, 2 Nanyang Technological University  \n{liujlong27, [wanghb69](wanghb69}@mail2.sysu.edu.cn)[}](wanghb69}@mail2.sysu.edu.cn)[@mail2.sysu.edu.cn](wanghb69}@mail2.sysu.edu.cn), [luoweiqi@mail.sysu.edu.cn](luoweiqi@mail.sysu.edu.cn), [jiaxiaojunqaq@gmail.com](jiaxiaojunqaq@gmail.com)  \narXiv :2607 .0048 1v 1 [ cs .CR] 1 Jul 2026  \nAbstract—Jailbreak attacks remain a critical threat to the safe deployment of large language models (LLMs). While prior work has primarily studied attacks and defenses atthe prompt level, we show that this prompt-centric paradigm overlooks a structural vulnerability in stateful, function-calling environments. In such applications, developer-defined schemas, structured arguments, and untrusted tool outputs are interleaved into a single shared model context. This architecture expands the attack surface by blurring the boundary between trusted control logic and untrusted data, allowing adversarial intent to be distributed across a multi-turn execution path. We exploit this architectural flaw through SMT, a black-box attack framework based on Simulated Moderation Traces. Departing from purely prompt-based interactions, SMT constructs a multi-turn trajectory that simulates a legitimate moderationauditing workflow. Within this trajectory, a fabricated moderation frame leverages red-team testing as a pretext toelicit harmful generations. The subsequent validation feedback treats safety refusals as execution failures, prompting refinements that gradually weaken the model’s safety constraints and ultimately trigger harmful outputs. Extensive empirical evaluations on prominent commercial LLMs from five different providers across two standardized safety benchmarks show that SMT consistently achieves the highest average attack success rate and HarmScore while requiring a near-minimal number of queries, substantially outperforming existing baselines. These findings demonstrate that prompt-level sanitization alone is fundamentally insufficient for defending tool-enabled LLM systems and highlight the urgent need for contextaware validation across schemas, arguments, tool outputs, and accumulated conversation state. The code is available at [https://github.com/liujlong27/SMT](https://github.com/liujlong27/SMT).  \n1. Introduction  \nLarge language models (LLMs) are increasingly deployed as software assistants [1],[2], autonomous agents [3], and tool-augmented applications’ core reasoning and con  \ntrol components [4] . To reduce the risk of generating il-+Equal contribution. * Corresponding author.  \nlegal, toxic, or harmful content [5], [6], model providers employ post-training alignment techniques, such as supervised safety fine-tuning and reinforcement learning from human feedback (RLHF) [7], [8] . Nevertheless, safety-aligned LLMs remain vulnerable to jailbreak attacks [9], [10], in which adversaries use carefully crafted prompts [11], deceptive contextual framing [12], or iterative inducement strategies [13], [14] to circumvent refusal mechanisms and elicit prohibited content [15] . Understanding such failures is crucial as LLMs move beyond standalone chat interfaces to production workflows that maintain interaction histories, invoke external functions, and process multi-source data.  \nMost existing jailbreak research adopts a prompt-centric threat model, treating the user-visible prompt as the primary attack surface. Single-turn methods manipulate the wording, representation, or structure of harmful requests. For example, ArtPrompt [17] obfuscates sensitive terms with ASCII-art representations, while CC-BOS [14] optimizes black-box jailbreak prompts using classical Chinese contexts and multi-dimensional fruit fly optimization, Do Anything Now [","cbCaicUwzPlHW6Ov","https://ap.wps.com/l/cbCaicUwzPlHW6Ov","pdf",1567526,4,1,19,"English","en",105,"# Abstract\n# Introduction\n## Motivation and threat model\n## Limitations of prompt-centric attacks\n## Function-calling architectural vulnerability","[{\"question\":\"What key vulnerability does the paper highlight in function-calling LLM systems?\",\"answer\":\"It shows that shared context windows interleave trusted schemas, structured arguments, tool outputs, and prior states, blurring the boundary between control logic and untrusted data across multiple turns.\"},{\"question\":\"How does the SMT method generate an attack trajectory?\",\"answer\":\"SMT constructs a multi-turn simulated moderation auditing workflow and uses a fabricated moderation frame to obtain harmful generations under the pretext of red-team testing.\"},{\"question\":\"What do the evaluation results conclude about prompt-level defenses?\",\"answer\":\"Across multiple commercial LLMs and two safety benchmarks, SMT achieves the highest average attack success and HarmScore with near-minimal queries, indicating prompt-level sanitization alone cannot reliably defend tool-enabled LLM systems.\"}]",1784188434,48,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"beyond-the-prompt-jailbreaking-function-calling-llms-via-simulated-moderation-traces","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/beyond-the-prompt-jailbreaking-function-calling-llms-via-simulated-moderation-traces/83495/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What key vulnerability does the paper highlight in function-calling LLM systems?","Question",{"text":75,"@type":76},"It shows that shared context windows interleave trusted schemas, structured arguments, tool outputs, and prior states, blurring the boundary between control logic and untrusted data across multiple turns.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the SMT method generate an attack trajectory?",{"text":80,"@type":76},"SMT constructs a multi-turn simulated moderation auditing workflow and uses a fabricated moderation frame to obtain harmful generations under the pretext of red-team testing.",{"name":82,"@type":73,"acceptedAnswer":83},"What do the evaluation results conclude about prompt-level defenses?",{"text":84,"@type":76},"Across multiple commercial LLMs and two safety benchmarks, SMT achieves the highest average attack success and HarmScore with near-minimal queries, indicating prompt-level sanitization alone cannot reliably defend tool-enabled LLM systems.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},"General","general"]