[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84898-en":3,"doc-seo-84898-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},84898,8796095461610,"Oliver","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","AEGIS A Mechanism Guided Defense against Visual Synonym Jailbreaks in Text-to-Image Models","Text-to-image diffusion models deliver high visual fidelity yet remain vulnerable to safety violations when adversaries induce illicit content. Existing alignment methods—such as input sanitization or pruning of structural features—focus mainly on unsafe concepts explicitly exposed during filtering or editing, leaving a blind spot for visual synonym attacks where benign prompts trigger prohibited imagery via implicit visual-semantic associations. This work introduces a mechanistic view showing convergence to sparse semantic-injecting attention heads, then proposes AEGIS to adaptively steer only identified vulnerable heads for improved safety and utility.","AEGIS: A Mechanism-Guided Defense against Visual Synonym Jailbreaks in Text-to-Image Models  \nYuanmin Huang, Zhenfei Zhang, Mi Zhang, Geng Hong, Qinqin He, Jialing Tao, Hui Xue, and Min Yang  \narXiv :2607 .06 120v 1 [ cs .CV] 7 Jul 2026  \nAbstract—Text-to-image diffusion models have achieved high visual fidelity and broad adoption, but remain vulnerable to safety violations when adversaries exploit them to synthesize illicit content. Existing alignment paradigms, from input sanitization to structural feature pruning, are largely organized around unsafe concepts explicitly exposed during filtering, editing, or localization. This leaves a blind spot for visual synonym attacks (VSA), a jailbreak where benign-looking prompts elicit prohibited imagery through implicit visual associations. As a result, current defenses face a safety-utility dilemma: they may either under-mitigate VSA threats or over-suppress visually similar benign concepts. The core challenge is that VSA hides the unsafe target at the textual surface while revealing it through generation-time visualsemantic convergence. In this work, we therefore shift from static suppression of pre-specified unsafe concepts to dynamic tracing of how unsafe semantics emerge during generation. Our mechanistic analysis shows that VSA and explicit unsafe prompts converge through sparse semantic-injecting attention heads, which serve as inference-time bottlenecks for prohibited visual semantics. Based on this insight, we propose AEGIS (Adaptive Evasion Guard via Identification and Steering), an inference-time defense that applies similarity-aware repulsion only at the identified vulnerable heads. Evaluated against 16 baselines, AEGIS improves both safety and utility. On SD 1.4, it reduces ASR to 0 .00/0 .03 for indomain violence/nudity VSA and achieves ASRs ≤ 0.09 on outof-domain explicit and adversarial attacks. It preserves benign fidelity, avoids suppressing hard-negative concepts, and transfers to SD 2.1 and FLUX.1 after re-identifying the critical heads for each backbone.  \nWarning: This paper contains sensitive content that may be disturbing or offensive to readers.  \nIndex Terms—Text-to-Image Diffusion Models, Visual Synonym Attacks, Safety Alignment, Mechanistic Interpretability.  \nI. INTRODUCTION  \nText-to-image (T2I) diffusion models have been widely deployed in real-world applications, enabling the generation of high-fidelity visual content from natural language descriptions [1]–[3] . Despite their impressive capability, the openended nature of these models introduces severe security and safety risks in production environments. Without effective  \nYuanmin Huang, Zhenfei Zhang, Mi Zhang, Geng Hong, and Min Yang are with Fudan University, Shanghai, China (e-mail: [yuanminhuang23@m.fudan.edu.cn](yuanminhuang23@m.fudan.edu.cn); [zhangzf24@m.fudan.edu.cn](zhangzf24@m.fudan.edu.cn);  \nmi [zhang@fudan.edu.cn](zhang@fudan.edu.cn); [ghong@fudan.edu.cn](ghong@fudan.edu.cn); [m](m yang@fudan.edu.cn)[ ](m yang@fudan.edu.cn)[yang@fudan.edu.cn](m yang@fudan.edu.cn)).  \nQinqin He, Jialing Tao, and Hui Xue are with Alibaba Group, Hangzhou, China (e-mail: [heqinqin.hqq@alibaba-inc.com](heqinqin.hqq@alibaba-inc.com); [jialing.tjl@alibaba-inc.com](jialing.tjl@alibaba-inc.com); [hui.xueh@alibaba-inc.com](hui.xueh@alibaba-inc.com)) .  \nMin Yang is a faculty of Shanghai Pudong Research Institute of Cryptology, and Engineering Research Center of Cyber Security Auditing and Monitoring, Ministry of Education, China.  \n\n| (a) Input\u003Cbr>[Explicit]\u003Cbr>Blood\u003Cbr>[Synonym]\u003Cbr>Spilled Dark\u003Cbr>Red Paint\u003Cbr> |  |  | [Benign]\u003Cbr>Ketchup\u003Cbr> |  (b) Under-Mitigation\u003Cbr>\u003Cbr>Aligned  Bypassed Preserved  |  |\n| --- | --- | --- | --- | --- | --- |\n|  (c) Over-Mitigation |  |  |  |  |  |\n|  |  |  |  |  |  |\n\nFig. 1: (a) Unprotected generation of explicit, visual synonym, and benign prompts. Generation of visual synonyms resembles explicit prompts, despite their semantic orthogonality in text space. (b) Under-mitigati","cbCaiqlwTBSyjPRD","https://ap.wps.com/l/cbCaiqlwTBSyjPRD","pdf",45604782,4,1,23,"English","en",105,"# Introduction\n## Evolving Attacks\n## Safety Alignment Paradigms","[{\"question\":\"What vulnerability does AEGIS target in text-to-image models?\",\"answer\":\"AEGIS targets visual synonym jailbreaks, where prompts that look benign in text can still produce prohibited imagery through implicit visual-semantic associations.\"},{\"question\":\"How do existing defenses tend to fail against visual synonym attacks?\",\"answer\":\"They often suppress only unsafe concepts explicitly exposed to text-centric filtering or structural pruning, which creates a safety-utility trade-off that can under-mitigate or over-suppress visually similar benign concepts.\"},{\"question\":\"What is the main idea behind AEGIS and how does it improve safety?\",\"answer\":\"AEGIS performs inference-time identification and steering of sparse semantic-injecting attention heads that act as bottlenecks for prohibited visual semantics, applying similarity-aware repulsion only at those vulnerable heads to reduce attack success while preserving benign fidelity.\"}]",1784199193,58,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"aegis-a-mechanism-guided-defense-against-visual-synonym-jailbreaks-in-text-to-image-models","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/aegis-a-mechanism-guided-defense-against-visual-synonym-jailbreaks-in-text-to-image-models/84898/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-23","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What vulnerability does AEGIS target in text-to-image models?","Question",{"text":75,"@type":76},"AEGIS targets visual synonym jailbreaks, where prompts that look benign in text can still produce prohibited imagery through implicit visual-semantic associations.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How do existing defenses tend to fail against visual synonym attacks?",{"text":80,"@type":76},"They often suppress only unsafe concepts explicitly exposed to text-centric filtering or structural pruning, which creates a safety-utility trade-off that can under-mitigate or over-suppress visually similar benign concepts.",{"name":82,"@type":73,"acceptedAnswer":83},"What is the main idea behind AEGIS and how does it improve safety?",{"text":84,"@type":76},"AEGIS performs inference-time identification and steering of sparse semantic-injecting attention heads that act as bottlenecks for prohibited visual semantics, applying similarity-aware repulsion only at those vulnerable heads to reduce attack success while preserving benign fidelity.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]