[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82507-en":3,"doc-seo-82507-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},82507,687197100911,"Himbo","https://ap-avatar.wpscdn.com/avatar/a000239b6f1da00475?x-image-process=image/resize,m_fixed,w_180,h_180&k=1782698725881665579",8,"Research & Report","The Illusion of High Utility in Safety Alignment of Text-to-Image Diffusion Models","Safety alignment for text-to-image (T2I) diffusion models suppresses harmful generations while preserving image utility on benign prompts. Prior work often reports high safety alongside high utility, but the impression is driven by coarse global metrics such as FID and CLIPScore that overlook fine-grained semantic correctness. Structured evaluation breaks this illusion: TIFA shows safety-aligned models suffer notable semantic fidelity drops, including errors in object counts, attributes, and relationships. The paper identifies semantic collapse in prompt embedding geometry and introduces StructureAware Geometric Regularization (SAGE) to preserve embedding spread and inter-prompt relational structure, restoring structured utility while retaining strong safety and competitive coarse scores.","arXiv :2607 .00402v 1 [ cs .CV] 1 Jul 2026  \nThe Illusion of High Utility in Safety Alignment of Text-to-Image Diffusion Models  \nAdeel Yousaf, Soumik Ghosh, James Beetham, Amrit Singh Bedi, and Mubarak Shah  \nInstitute of Artificial Intelligence  \nUniversity of Central Florida, Orlando, United States {adeel.yousaf,soumik.ghosh,james.beetham,[amritbedi}@ucf.edu](amritbedi}@ucf.edu) , [shah@crcv.ucf.edu](shah@crcv.ucf.edu)  \nAbstract. Safety alignment of text-to-image (T2I) diffusion models aims to suppress harmful generations while preserving utility on benign prompts.  \nRecent methods often appear to deliver high safety with high utility, but this conclusion rests largely on coarse global utility metrics (e.g. , FID, CLIPScore) that are insensitive to fine-grained semantic correctness, creating an illusion of high utility. We show that when utility is measured with structured evaluation, this illusion breaks: on TIFA (Text-to-Image Faithfulness evaluation with Question Answering), safety-aligned models suffer substantial drops in semantic fidelity, including failures in object counts, attributes, and relationships. To diagnose the source of this gap, we analyze the text-encoder prompt embedding space and uncover semantic collapse, a contraction of embedding spread coupled with distortion of inter-prompt similarity structure, which strongly correlates with structured utility loss. Guided by this insight, we propose StructureAware Geometric Regularization (SAGE) 1 , a safety alignment objective that explicitly preserves embedding spread and inter-prompt relational structure during adaptation. Our method restores structured utility (TIFA +5.0% over prior state-of-the-art) while maintaining strong safety performance and competitive coarse-grained utility scores.  \nKeywords: Text-to-Image · Safety Alignment · Embedding Analysis  \n1 Introduction  \nText-to-image (T2I) diffusion models such as stable diffusion (SD) [24] and DALL-E [20] can generate highly realistic images from natural language prompts, but their training on large web-scale datasets also exposes them to unsafe concepts, enabling the generation of Not-Safe-For-Work (NSFW) content [25, 38] . As these models become widely accessible, mitigating unsafe generations has become a critical requirement for responsible deployment. At the same time, safety modifications must preserve the model’s core capability: generating images that  \n1 [https://adeelyousaf.github.io/SAGE_ECCV26_Project_Page/](https://adeelyousaf.github.io/SAGE_ECCV26_Project_Page/)  \n2 A. Yousaf et al.  \nPrompt  \nA blue bird with a bright yellow beak. A shelf with three vases: red, green, and blue.  \nSD v1.4 Base Model Safe Model Base Model Safe Model  \nGenerated Image  \nCurrent Coarse Metric (CLIPScore)  \nFine-grained Metric (TIFA)  \n\n|  |  |  |  |\n| --- | --- | --- | --- |\n\n\n| \u003Cbr>score = 0.85 |  |  |  |\n| --- | --- | --- | --- |\n\nFig. 1: The illusion of high utility under coarse evaluation. Comparison between the base model and a safe model (unlearned) across fine-grained prompts. While the safe model fails to generate specific attributes (e.g., the yellow beak or the correct vase colors/count ), the standard CLIPScore provides misleadingly higher scores for the incorrect images (✗) . In contrast, the fine-grained metric, TIFA, accurately captures the utility degradation (✓), properly penalizing the safe model for failing to satisfy the detailed visual requirements of the prompt.  \nfollow the prompt. If enforcing safety significantly degrades model capability, aligned models may become less attractive than their unrestricted counterparts.  \nRecently, several works have proposed safety alignment methods for T2I models to suppress harmful generations [1, 15 , 21 , 25 , 28 , 36–38] . These approaches typically evaluate safety using attack success rates (ASR), while utility is assessed using coarse metrics such as Fréchet Inception Distance (FID) [9] and CLIPScore. In early works, improving safety often came a","cbCaicX6QN5zYxJy","https://ap.wps.com/l/cbCaicX6QN5zYxJy","pdf",9017501,4,1,37,"English","en",105,"# Abstract\n# Introduction\n## An illusion of high utility\n## Our diagnosis: semantic collapse","[{\"question\":\"Why do safety-aligned T2I models sometimes appear to have high utility?\",\"answer\":\"Because common utility reports rely on coarse global metrics like FID and CLIPScore, which are insensitive to fine-grained compositional and semantic correctness.\"},{\"question\":\"What does TIFA reveal about the utility of safety-aligned models?\",\"answer\":\"Structured evaluation with TIFA shows substantial drops in semantic fidelity, including failures in object counts, attributes, and relationships, even when coarse metrics look competitive.\"},{\"question\":\"What is the main cause proposed for the structured utility loss?\",\"answer\":\"The paper attributes the gap to semantic collapse in the text-encoder prompt embedding space, characterized by contraction of embedding spread and distortion of inter-prompt similarity structure.\"}]",1784181015,93,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"the-illusion-of-high-utility-in-safety-alignment-of-text-to-image-diffusion-models","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/the-illusion-of-high-utility-in-safety-alignment-of-text-to-image-diffusion-models/82507/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-23","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why do safety-aligned T2I models sometimes appear to have high utility?","Question",{"text":75,"@type":76},"Because common utility reports rely on coarse global metrics like FID and CLIPScore, which are insensitive to fine-grained compositional and semantic correctness.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What does TIFA reveal about the utility of safety-aligned models?",{"text":80,"@type":76},"Structured evaluation with TIFA shows substantial drops in semantic fidelity, including failures in object counts, attributes, and relationships, even when coarse metrics look competitive.",{"name":82,"@type":73,"acceptedAnswer":83},"What is the main cause proposed for the structured utility loss?",{"text":84,"@type":76},"The paper attributes the gap to semantic collapse in the text-encoder prompt embedding space, characterized by contraction of embedding spread and distortion of inter-prompt similarity structure.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]