[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86385-en":3,"doc-seo-86385-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86385,549758146520,"Patrick","https://ap-avatar.wpscdn.com/avatar/80002397d8c0411e94?_k=1775819394049821470",8,"Research & Report","SPQR: A Multi-Dimensional Benchmark for Safety Alignment under Benign Model Adaptation","Text-to-image diffusion models can generate copyrighted, unsafe, or private content, while safety alignment methods are usually evaluated only at release time. This work studies whether safety persists after benign, post-deployment fine-tuning such as LoRA personalization and style/domain adapters. Experiments reveal frequent safety breakdowns despite quality gains. To address this, SPQR introduces a unified Safety–Prompt adherence–Quality–Robustness benchmark, summarized by a harmonic mean score, enabling stable, reproducible evaluation and clearer detection of regressions.","SPQR: A Multi-Dimensional Benchmark for Safety Alignment under Benign Model Adaptation  \nMohammed Talha Alam 1, Nada Saadi 1, Fahad Shamshad 1, Nils Lukas 1, Karthik Nandakumar 1 ,3, Fakhri Karray 1 ,2, and Samuele Poppi 1  \n1 Mohamed bin Zayed University of Artificial Intelligence (MBZUAI), UAE  \n2 University of Waterloo, Canada  \n3 Michigan State University, USA  \n{mohammed.alam, nada.saadi, fahad.shamshad, nils.lukas, karthik.nandakumar, fakhri.karray, [samuele.poppi}@mbzuai.ac.ae](samuele.poppi}@mbzuai.ac.ae)  \narXiv :2511 . 19558v2 [ cs .CR] 11 Jul 2026  \n(a) Safety Failure  \nAnimage of a [harmful  \nobject] Benign  \nAnimage of a  \n[harmful object]  \nFine-Tuning  \n| Safe Generation\u003Cbr> |  | Unsafe Generation\u003Cbr> |\n| --- | --- | --- |\n|  |  |  |\n| Quality improves while safety collapses Failures are invisible to standard metrics! |  |  |\n\nHarmonic Mean Score  \nS: Safety  \nUnsafe content unlearning  \nQ: Quality  \nVisual ﬁdelity (FID)  \nP: Prompt adherence  \nText-image alignment  \nR: Robustness  \nPost-FT stability  \n(c) Method Comparison  \nFig. 1: Illustration of the SPQR benchmark. (Left) Example of a benign fine-tuning (BFT) causing safety regression: before BFT, Stable Diffusion produces a safe image, while after BFT, the same prompt yields a harmful one. (Center) SPQR evaluates models along four axes—Safety (S), Prompt Adherence (P), Quality (Q), and Robustness (R)—and aggregates them into a single harmonic mean score. (Right) Representative comparison showing how different safety-alignment methods vary across dimensions, highlighting strong pre-adaptation safety but weak robustness after benign fine-tuning.  \nAbstract. Text-to-image diffusion models can emit copyrighted, unsafe, or private content. Safety alignment aims to suppress specific concepts, yet evaluations seldom test whether safety persists under benign downstream fine-tuning routinely applied after deployment (e.g., LoRA personalization, style/domain adapters) . We study the stability of current safety methods under benign fine-tuning and observe frequent breakdowns.  \nAs true safety alignment must withstand even benign post-deployment adaptations, we introduce the SPQR benchmark (Safety–Prompt adherence–Quality–Robustness) . SPQR is a single-scored metric that provides  \n a unified, reproducible framework to evaluate how well safety-aligned Code is available at [https://github.com/talha-alam/spqr](https://github.com/talha-alam/spqr).  \n2 M.T. Alam et al.  \ndiffusion models preserve safety, utility, and robustness under benign fine-tuning, by reporting a single leaderboard score to facilitate comparisons. We conduct multilingual, domain-specific, and out-of-distribution analyses, along with category-wise breakdowns, to identify when safety alignment fails after benign fine-tuning, ultimately showcasing SPQR as a concise yet comprehensive benchmark for T2I safety alignment techniques for T2I models.  \nWarning: This paper features illustrative examples that may involve explicit, sexual, or violent imagery and language, which could be sensitive for certain readers.  \n1 Introduction  \nThe widespread adoption of powerful open-source text-to-image models [35, 43] such as Stable Diffusion (SD) [35] has democratized content creation, but also created a need for reliable safety and control [35, 36] . These models can memorize and regenerate harmful [3] or copyrighted [4, 38] training content, and unsafe prompts can induce unsafe outputs via cross-attention conditioning. Moreover, even safe prompts can yield unsafe generations if model priors or dataset biases activate correlated directions in the CLIP embedding space [21, 32] . The practical question is not only how to prevent unsafe outputs at release time, but how to ensure that prevention persists throughout the model’s lifecycle. From the early Safety Checker [35] for Stable Diffusion, many methods have emerged [6] . Despite different mechanisms, their shared objective is to reduce the probability of unsafe genera","cbCaintcfoLliinI","https://ap.wps.com/l/cbCaintcfoLliinI","pdf",28319569,7,1,34,"English","en",105,"# Introduction\n## Safety alignment and evaluation gaps\n## Robustness to attacks and benign fine-tuning\n## SPQR benchmark overview","[{\"question\":\"What problem does SPQR address in safety alignment for text-to-image diffusion models?\",\"answer\":\"SPQR addresses the lack of evaluation for whether safety alignment remains effective after benign post-deployment fine-tuning, which can cause safety regressions.\"},{\"question\":\"What is “benign model adaptation” in this paper?\",\"answer\":\"It refers to downstream fine-tuning such as LoRA personalization, style adapters, or domain adapters applied on strictly harmless data, which can unintentionally weaken or reverse safety alignment.\"},{\"question\":\"How does SPQR measure safety alignment?\",\"answer\":\"SPQR evaluates models across four dimensions—Safety, Prompt adherence, Quality, and Robustness—and aggregates them into a single harmonic mean score for easy comparison.\"}]",1784211411,86,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"spqr-a-multi-dimensional-benchmark-for-safety-alignment-under-benign-model-adaptation","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/spqr-a-multi-dimensional-benchmark-for-safety-alignment-under-benign-model-adaptation/86385/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What problem does SPQR address in safety alignment for text-to-image diffusion models?","Question",{"text":76,"@type":77},"SPQR addresses the lack of evaluation for whether safety alignment remains effective after benign post-deployment fine-tuning, which can cause safety regressions.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"What is “benign model adaptation” in this paper?",{"text":81,"@type":77},"It refers to downstream fine-tuning such as LoRA personalization, style adapters, or domain adapters applied on strictly harmless data, which can unintentionally weaken or reverse safety alignment.",{"name":83,"@type":74,"acceptedAnswer":84},"How does SPQR measure safety alignment?",{"text":85,"@type":77},"SPQR evaluates models across four dimensions—Safety, Prompt adherence, Quality, and Robustness—and aggregates them into a single harmonic mean score for easy comparison.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":107,"slug":138},19,"General","general"]