[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83455-en":3,"doc-seo-83455-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83455,1099513958762,"Logic","https://ap-avatar.wpscdn.com/avatar/1000023916a998db790?x-image-process=image/resize,m_fixed,w_180,h_180&k=1784791008015729253",8,"Research & Report","DriftScope: Measuring The Hidden Effects of Diffusion Model Adaptation","Adapting pre-trained text-to-image diffusion models is often judged only by intended outcomes such as learning new visual concepts or removing unwanted ones. DriftScope argues this evaluation is incomplete by showing that weight-level adaptation systematically damages semantically unrelated concepts. Sparse autoencoder analysis and zero-shot classification reveal either aggregate metrics remain flat while per-class zero-shot accuracy drops up to 18.9 points, or models become nearly unusable. DriftScope provides prompt-level token rankings of concept drift.","DriftScope: Measuring The Hidden Effects of Diffusion Model Adaptation  \narXiv :2607 .00183v1 [ cs .CV] 30 Jun 2026  \nHéctor Laria∗ 1 ,2, Yiping Han∗ 1 ,2, Julian D. Santamaria 1 ,2, Kai Wang3 ,4†, Bogdan Raducanu 1 ,2, Joost van de Weijer 1 ,2, and Alexandra Gomez-Villa 1 ,2  \n1 Computer Vision Center, Barcelona, Spain  \n2 Universitat Autonoma de Barcelona, Barcelona, Spain  \n3 Program of Computer Science, City University of Hong Kong (Dongguan), China  \n4 City University of Hong Kong, HK SAR, China {hlaria, yhan, jsantamaria, bogdan, joost, [agomezvi}@cvc.uab.es](agomezvi}@cvc.uab.es) ,  \n[kai.wang@cityu-dg.edu.cn](kai.wang@cityu-dg.edu.cn)  \nAbstract. Adapting pre-trained text-to-image diffusion models, whether to learn new visual concepts or erase unwanted ones, is routinely evaluated on its intended effects alone. We argue this framing is incomplete. Through sparse autoencoder analysis and zero-shot classification, we demonstrate that adaptation systematically damages semantically unrelated concepts in ways that aggregate metrics structurally cannot surface: when damage is severe enough for FID and KID to respond, the model is already nearly unusable; when the model remains functional, FID and KID stay flat while specific classes silently suffer worst-case zero-shot accuracy drops of up to 18.9 points and concept-level distributions shift dramatically. This pattern appears at both ends of the adaptation spectrum (concept customization and concept unlearning), suggesting it is a systematic consequence of weight-level modification rather than an artifact of any particular method. To surface this hidden drift before deployment, we introduce DriftScope, a prompt-level diagnostic tool that takes any two model checkpoints and returns a ranked list of tokens whose visual concepts have shifted most between them.  \nDriftScope optimizes a soft prompt to attribute drift at the token level without requiring access to real data or model internals. The result isan interpretable, concept-level audit that aggregate evaluation cannot provide.  \nKeywords: Diffusion models · Model evaluation · Model diffing  \n1 Introduction  \nAdapting foundation models to specific needs has become a central challenge in modern AI deployment, and weight-level adaptation is among the most widely  \n∗ Equal Contribution.  \n† Corresponding Author.  \nProject page: [https://hyping111.github.io/DriftScope/](https://hyping111.github.io/DriftScope/)  \n2 H. Laria et al.  \nadopted solutions. Whether customizing text-to-image models to new visual concepts [26] or removing unsafe content through unlearning [10], practitioners routinely modify pre-trained weights to serve specific needs. These operations are treated as benign optimizations, evaluated mostly on their intended effects: Does the customized model generate the new concept? Does the unlearned model avoid harmful content?  \nWe argue this framing is incomplete. While adaptation achieves its proximal goals, it simultaneously degrades the model in ways current evaluation misses [17] . Standard metrics (such as FID, KID, prompt-specific evaluations) focus on aggregate quality and target-task performance, missing fine-grained distributional shifts. As a result, adapted models may lose entire categories of concepts and exhibit dramatic perceptual shifts on unrelated tasks, all while passing standard benchmarks. These “hidden costs” arise across opposite ends of the adaptation spectrum: concept customization, which adds a new concept, and concept unlearning, which removes one, suggesting this is a systematic consequence of weight-level adaptation, not an artifact of any particular method.  \nThe core problem is granularity. Standard metrics measure aggregate quality; they are structurally blind to concentrated damage on specific concepts. What practitioners need is not another aggregate score, but a tool that answers a precise question: Given this prompt, which words have their generated concepts most affected by adapta","cbCaifA9n4KMFZ4j","https://ap.wps.com/l/cbCaifA9n4KMFZ4j","pdf",4206618,5,1,22,"English","en",105,"# Introduction\n## Problem framing: hidden costs of weight-level adaptation\n## Evidence and analyses: sparse autoencoders and zero-shot classification\n## DriftScope: prompt-level diagnostic via cross-attention divergence","[{\"question\":\"Why do FID and KID fail to capture the effects of diffusion model adaptation?\",\"answer\":\"They focus on aggregate quality, so concept-level collateral damage can accumulate without changing these scores. When the damage becomes severe enough to move FID/KID, the model is already nearly unusable.\"},{\"question\":\"What evidence supports the claim that adaptation harms semantically unrelated concepts?\",\"answer\":\"Sparse autoencoder analysis surfaces concept-level shifts in the generative distribution, and zero-shot classification confirms the shifts appear in the model’s latent space.\"},{\"question\":\"How does DriftScope identify which tokens are most affected by adaptation?\",\"answer\":\"DriftScope uses a prompt generated by the base model and feeds it through both checkpoints, then maximizes divergence of cross-attention maps. Tokens with the largest cross-attention differences are ranked as the most drifted visual concepts.\"}]",1784188073,55,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"driftscope-measuring-the-hidden-effects-of-diffusion-model-adaptation","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/driftscope-measuring-the-hidden-effects-of-diffusion-model-adaptation/83455/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"Why do FID and KID fail to capture the effects of diffusion model adaptation?","Question",{"text":76,"@type":77},"They focus on aggregate quality, so concept-level collateral damage can accumulate without changing these scores. When the damage becomes severe enough to move FID/KID, the model is already nearly unusable.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"What evidence supports the claim that adaptation harms semantically unrelated concepts?",{"text":81,"@type":77},"Sparse autoencoder analysis surfaces concept-level shifts in the generative distribution, and zero-shot classification confirms the shifts appear in the model’s latent space.",{"name":83,"@type":74,"acceptedAnswer":84},"How does DriftScope identify which tokens are most affected by adaptation?",{"text":85,"@type":77},"DriftScope uses a prompt generated by the base model and feeds it through both checkpoints, then maximizes divergence of cross-attention maps. Tokens with the largest cross-attention differences are ranked as the most drifted visual concepts.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":20,"slug":138},19,"General","general"]