[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83297-en":3,"doc-seo-83297-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83297,1374391974564,"Clementine","https://ap-avatar.wpscdn.com/avatar/14000253aa45c000a9e?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779874745381141002",8,"Research & Report","Open Models Open Risks Measuring Unsafe Generation in Text-to-Image Models In the Wild","Existing safety research on text-to-image (T2I) jailbreaks mostly relies on controlled in-the-lab experiments using a limited set of canonical models, leaving the real safety state of the fast-growing in-the-wild T2I ecosystem unclear. This gap is driven by detector-oriented metrics that target controlled evaluation and by unsafe risks that may stem from both adversarial prompting and unsafe release practices or risky derivatives. The study conducts a large-scale empirical jailbreak analysis across 200+ Hugging Face models.","Open Models, Open Risks: Measuring Unsafe Generation in Text-to-Image Models In the Wild  \nPeilin Han  \n[hanpeilin6788@gmail.com](hanpeilin6788@gmail.com)[ ](hanpeilin6788@gmail.com)Xidian University China  \nJingchun Zhang  \nXidian University China  \nYang Liu  \nXidian University China  \nTeng Li  \nXidian University China  \nYilong Yang  \nXidian University China  \nJianfeng Ma  \nXidian University China  \nZhuo Ma  \nXidian University China  \narXiv :2607 .07827v 1 [ cs .CR] 8 Jul 2026  \nAbstract  \nExisting safety studies on text-to-image (T2I) jailbreaks are largely conducted in controlled in-the-lab settings, typically on a small number of canonical models. As a result, the current safety status of the rapidly growing in-the-wild T2I ecosystem remains unclear. This uncertainty is amplified by two factors: existing detector-based metrics are designed for controlled evaluation, and in-the-wild risks may arise not only from adversarial prompting, but also from unsafe release practices and unsafe model derivatives.  \nIn this paper, we present a large-scale empirical study of in-thewild T2I safety through the lens of jailbreak. We first show that detector-only jailbreak metrics substantially overestimate practical risk over in the wild due to semantic drift and generation artifacts, and we introduce Advanced ASR to better capture semantically valid and visually plausible unsafe generation. Using this refined metric, we evaluate 200+ in-the-wild T2I models from Hugging Face under three representative jailbreak attacks. Our results show that many downstream models retain a non-trivial degree of safety even without explicit post-hoc safeguards, indicating that safety degradation in the wild is neither universal nor uniform. At the same time, we identify a set of high-risk models, including explicitly NSFW-oriented releases as well as seemingly benign models whose unsafe behavior is only exposed through systematic evaluation. We further trace these models to their release context and report high-risk cases to Hugging Face.  \nKeywords  \nText-to-Image Model, Jailbreak, Not Safe For Work, In the Wild  \n1 Introduction  \nWith the rapid advancement of diffusion-based architectures, multimodal generative models, particularly Text-to-Image (T2I) systems, have been widely adopted in real-world applications[5, 22, 30] . These models are capable of generating high-quality and visually coherent content, leading to a rapidly expanding user base. However, models developed in controlled laboratory environments (i.e., in-the-lab) are typically optimized for general-purpose objectives. Such designs are often insufficient to accommodate diverse and evolving user requirements. In practice, communities exhibit a  \nstrong demand for customization capabilities, including support for specific artistic styles and domain-specific generation tasks[3, 25] .  \nTo address this limitation, open model ecosystems have emerged on platforms such as Hugging Face and ModelScope [7, 21] . These platforms lower the barrier to model access and modification, enabling users to fine-tune and redistribute customized T2I models. Models deployed in the wild are often released with weakened, optional, or entirely removed safety mechanisms. However, the relaxation or removal of safety mechanisms exposes new attack surfaces. Among them, jailbreak-based manipulation[4, 9, 16, 28, 29] has emerged as an effective strategy to bypass content safeguards. By crafting specific prompts or conditioning inputs, adversaries can induce T2I models to generate unsafe or Not Safe for Work (NSFW) content that violate predefined safety policies.  \nOur work. In this paper, we systematically study the current safety status of in-the-wild T2I models through the following three research questions:  \n(1) RQ1: In-the-wild Jailbreak Metrics. What is the difference between lab and wild? What limitations arise from these existing evaluation metrics? How to accurately evaluate effective real-world risk?  \n(2) RQ2: S","cbCaioZAUZJ8LKPn","https://ap.wps.com/l/cbCaioZAUZJ8LKPn","pdf",3910416,5,1,17,"English","en",105,"# Introduction\n## In-the-Wild vs In-the-Lab Safety\n## Research Questions\n## Refined Evaluation Metrics\n## Large-Scale Evaluation Setup\n## Findings and High-Risk Identification","[{\"question\":\"Why are current text-to-image jailbreak safety results unclear in real-world (in-the-wild) settings?\",\"answer\":\"Most studies use controlled in-the-lab evaluations with a small set of canonical models, and detector-based metrics are designed for those controlled conditions. Real-world risks can also emerge from unsafe release practices and unsafe model derivatives, not only adversarial prompts.\"},{\"question\":\"What problem do detector-only jailbreak metrics have when measuring in-the-wild risk?\",\"answer\":\"Detector-only metrics substantially overestimate practical risk because they can be unreliable in the wild due to semantic drift and generation artifacts. The study notes that detectors may flag images as unsafe even when they do not satisfy real-world NSFW objectives.\"},{\"question\":\"How does the paper evaluate in-the-wild safety and what does it conclude about model safety variation?\",\"answer\":\"It introduces an advanced ASR-style refined metric (AASR) to better capture semantically valid and visually plausible unsafe generation, then evaluates 200+ in-the-wild T2I models with three representative jailbreak attacks. Results indicate safety degradation is neither universal nor uniform: many downstream models retain non-trivial safety, while high-risk models are identified and traced for reporting.\"}]",1784186577,43,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"open-models-open-risks-measuring-unsafe-generation-in-text-to-image-models-in-the-wild","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/open-models-open-risks-measuring-unsafe-generation-in-text-to-image-models-in-the-wild/83297/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"Why are current text-to-image jailbreak safety results unclear in real-world (in-the-wild) settings?","Question",{"text":76,"@type":77},"Most studies use controlled in-the-lab evaluations with a small set of canonical models, and detector-based metrics are designed for those controlled conditions. Real-world risks can also emerge from unsafe release practices and unsafe model derivatives, not only adversarial prompts.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"What problem do detector-only jailbreak metrics have when measuring in-the-wild risk?",{"text":81,"@type":77},"Detector-only metrics substantially overestimate practical risk because they can be unreliable in the wild due to semantic drift and generation artifacts. The study notes that detectors may flag images as unsafe even when they do not satisfy real-world NSFW objectives.",{"name":83,"@type":74,"acceptedAnswer":84},"How does the paper evaluate in-the-wild safety and what does it conclude about model safety variation?",{"text":85,"@type":77},"It introduces an advanced ASR-style refined metric (AASR) to better capture semantically valid and visually plausible unsafe generation, then evaluates 200+ in-the-wild T2I models with three representative jailbreak attacks. Results indicate safety degradation is neither universal nor uniform: many downstream models retain non-trivial safety, while high-risk models are identified and traced for reporting.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":20,"slug":138},19,"General","general"]