[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83200-en":3,"doc-seo-83200-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83200,1374391974468,"Eden","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","Stage-Aware Adaptation and Distribution Calibration for Subject-Driven Personalized Text-to-Image Generation","Subject-driven personalized text-to-image generation adapts a pretrained diffusion model using only a few reference images to learn a specific subject identity, follow novel text prompts, and keep sample diversity. Existing low-rank personalization and adapter methods fail to separate capacity needs across diffusion denoising stages, and inference-time identity-only candidate selection can shrink visual diversity. The framework combines SPaRa stage-aware low-rank training with DCAL distribution-calibrated candidate selection. On SDXL and DreamBooth 30-subject protocols, DCAL improves multiple identity- and text-related metrics while exposing diversity tradeoffs, guiding evaluation beyond identity similarity alone.","Stage-Aware Adaptation and Distribution Calibration for Subject-Driven Personalized Text-to-Image Generation  \nWenyan Xu 1 ,∗ Alizer Wong2 ,3 ,∗  \n1 School of Computer Science, Guangdong University of Technology  \n2 School of Computer Science, Peking University 3ManXis  \n[3223004777@mail2.gdut.edu.com](3223004777@mail2.gdut.edu.com) [aliiiiezer@gmail.com](aliiiiezer@gmail.com)  \n∗Equal contribution  \narXiv :2607 .07 173v 1 [ cs .CV] 8 Jul 2026  \nAbstract  \nSubject-driven personalized text-to-image generation requires a pretrained diffusion model to acquire a specific subject from a few reference images while preserving subject identity, following novel text prompts, and maintaining sample diversity. Existing optimization-based methods instantiate subject adaptation through full fine-tuning, textual embedding optimization, or low-rank parameter updates; PaRa further constrains personalization from the perspective of parameter rank reduction. However, a uniform low-rank constraint or a uniform adapter strength cannot explicitly distinguish the capacity requirements of different denoising stages. Moreover, inference-time candidate selection driven mainly by identity similarity may compress the selected samples in the visual representation space. We decompose the problem into two complementary components: SPaRa denotes training-side stage-aware low-rank adaptation, DCAL denotes inference-side distribution-calibrated candidate selection, and SPaRa–DCAL denotes the combined framework. Theoretical analysis shows that timestep-dependent scaling controls the effective perturbation magnitude of a low-rank adapter, while identity-biased candidate selection restricts the radius of selected features around the reference center under explicit conditions. Auditable experiments under the SDXL and DreamBooth 30-subject protocol show that DCAL improves 1-LPIPS, CLIP-I, DINO-I, and CLIP-T on a fixed LoRA candidate pool, while revealing a clear tradeoff with CLIP/DINO pairwise diversity and pairwise LPIPS. These results indicate that personalized generation should be evaluated through identity consistency, text alignment, and representation diversity rather than identity metrics alone.  \n1 Introduction  \nText-to-image diffusion models have evolved from generalpurpose image synthesis systems into controllable genera-  \nFigure 1 . Motivation. Stage-agnostic adaptation and identity-only selection may both limit personalization quality.  \ntive infrastructure for real content production. Latent Diffusion Models and SDXL substantially improve visual fidelity, semantic compositionality, and cross-style transfer under natural-language conditions, which enables generative systems to support e-commerce material creation, intellectualproperty assets, personal digital avatars, advertising design, and interactive visual authoring [1, 2] . Practical generation targets, however, are rarely abstract categories alone. Many applications require a model to reproduce a concrete subject with stable identity, such as a particular pet, backpack, toy, commodity, or character. Subject-driven personalized text-to-image generation therefore asks a model to capture the fine-grained appearance that distinguishes one instance from other category members using only a few reference images, and then reuse those subject-specific attributes under new prompts describing unseen scenes, poses, materials, and styles [3, 4] . This ability determines whether diffusion models can move beyond open-domain synthesis toward user-level customization and asset-level reuse.  \nAs illustrated in Fig. 1, few-shot subject personalization is difficult because reference images provide the desired identity and incidental factors at the same time. Backgrounds, camera viewpoints, lighting conditions, occlusions, and local compositions are entangled with subject appearance in the small reference set. Excessive adaptation can absorb these incidental correlations, causing generated images to remain near th","cbCaijTYBCIzmMK2","https://ap.wps.com/l/cbCaijTYBCIzmMK2","pdf",5191754,2,1,16,"English","en",105,"# Abstract\n# Introduction","[{\"question\":\"What problem does stage-aware adaptation target in subject-driven personalized text-to-image generation?\",\"answer\":\"It addresses the mismatch between personalization capacity requirements across different diffusion denoising stages, which uniform low-rank constraints or adapter strengths cannot explicitly handle.\"},{\"question\":\"How does the proposed framework combine SPaRa and DCAL?\",\"answer\":\"SPaRa performs training-side stage-aware low-rank adaptation, while DCAL performs inference-side distribution-calibrated candidate selection; SPaRa–DCAL is evaluated as the combined framework.\"},{\"question\":\"What tradeoffs does DCAL introduce according to the experiments?\",\"answer\":\"DCAL improves identity- and text-alignment related metrics on a fixed LoRA candidate pool, while revealing a clear tradeoff with diversity measures such as CLIP/DINO pairwise diversity and pairwise LPIPS.\"}]",1784185912,40,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"stage-aware-adaptation-and-distribution-calibration-for-subject-driven-personalized-text-to-image-generation","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/stage-aware-adaptation-and-distribution-calibration-for-subject-driven-personalized-text-to-image-generation/83200/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-23","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does stage-aware adaptation target in subject-driven personalized text-to-image generation?","Question",{"text":75,"@type":76},"It addresses the mismatch between personalization capacity requirements across different diffusion denoising stages, which uniform low-rank constraints or adapter strengths cannot explicitly handle.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the proposed framework combine SPaRa and DCAL?",{"text":80,"@type":76},"SPaRa performs training-side stage-aware low-rank adaptation, while DCAL performs inference-side distribution-calibrated candidate selection; SPaRa–DCAL is evaluated as the combined framework.",{"name":82,"@type":73,"acceptedAnswer":83},"What tradeoffs does DCAL introduce according to the experiments?",{"text":84,"@type":76},"DCAL improves identity- and text-alignment related metrics on a fixed LoRA candidate pool, while revealing a clear tradeoff with diversity measures such as CLIP/DINO pairwise diversity and pairwise LPIPS.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,119,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":29,"slug":118},7,"Healthcare","healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]