[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86188-en":3,"doc-seo-86188-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86188,13056703019662,"Evangeline","https://ap-avatar.wpscdn.com/avatar/be000253a8e92610077?_k=1778726343310543188",8,"Research & Report","Structure-Detail Decoupled Autoregressive Generation for Fast and High-Fidelity Virtual Try-On","Virtual try-on (VTON) synthesizes a photorealistic person wearing a target garment, requiring faithful person preservation plus accurate garment deformation and detail synthesis. Diffusion-based VTON models combine these factors in compressed latent space but suffer from high inference latency and high-frequency detail loss from latent compression. Visual autoregressive generation offers faster inference yet lacks effective bi-conditional mechanisms for VTON. The document introduces VAR-VTON and STAR-VTON, decoupling structural synthesis in latent space from fine pixel-space detail recovery via a matching-informed refiner.","arXiv :2607 . 11233v1 [ cs .CV] 13 Jul 2026  \nSTRUCTURE-DETAIL DECOUPLED AUTOREGRESSIVE GENERATION FOR FAST AND HIGH-FIDELITY VIRTUAL TRY-ON  \nLu Yang1 , Xiaonan Hu1 , Yanan Li2 , Daqi Liu3 , Xiang Bai1 & Hao Lu1, ∗  \n1Huazhong University of Science and Technology, China  \n2Wuhan Institute of Technology, China  \n3Xiaomi EV, China  \n{lu yang1, [hlu](hlu}@hust.edu.cn)[}](hlu}@hust.edu.cn)[@hust.edu.cn](hlu}@hust.edu.cn)  \nABSTRACT  \nVirtual try-on (VTON) is a bi-conditional image generation problem that requires not only accurate person preservation but also faithful garment deformation and detail synthesis. Diffusion-based VTON methods can jointly model these factors in a compressed latent space, but suffer from high-frequency detail loss due to inherent latent compression, even with costly multi-step denoising. Recent visual autoregressive (VAR) models offer a promising alternative for high-quality generation with faster inference, yet remain unexplored for VTON due to the lack of effective bi-conditioning mechanisms. To bridge this gap, we first introduce VAR-VTON, a VAR-based VTON model that incorporates garment conditioning and structural guidance for efficient latent-space VTON. Despite its efficacy, latent-space generation still struggles to preserve fine-grained garment details. We argue that different VTON sub-tasks should be addressed in different representation spaces: structural synthesis such as garment warping and person layout is suited to the latent space, whereas fine-grained detail recovery should be tackled in the pixel space. Motivated by this insight, we further propose STARVTON, a Two-Stage AutoRegressive framework that builds upon VAR-VTON by decoupling latent-space structural synthesis from pixel-space detail recovery.  \nOur idea is to resort to a matching-informed refiner to establish dense correspondences between the stage-one generation and the source garment to directly map fine-grained pixel-space details. Extensive experiments show that STAR-VTON achieves an impressive efficiency–fidelity trade-off: VAR-VTON runs at least 4× faster than diffusion-based counterparts without degrading quality, and the pixelspace refiner effectively restores fine details and acts as a plug-and-play module that can benefit existing VTON approaches.  \n1 INTRODUCTION  \nGiven a person image and a garment image, virtual try-on (VTON) aims to synthesize aphotorealistic image in which the person is naturally and accurately dressed in the target garment (Han et al., 2018) . It is a typical bi-conditional generative problem, because a faithful VTON generation requires not only preserving the identity and pose of the person but also the fidelity of the garment.  \nAchieving high-fidelity VTON, however, is non-trivial. It needs to i) correctly deform the garment to align with the target body pose and  \n∗ Corresponding author.  \nFigure 1: Comparison with the state of the art in terms of detail fidelity and efficiency. Bubble size denotes computational cost (TFLOPs) . We evaluate detail preservation using CLIP-I, with higher values indicating better fidelity. Our solutions achieve better fidelity-efficiency trade-off.  \nFigure 2: Comparison between latent-space VTON and our structure-detail decoupled generation framework. (a) Existing VTON methods rely on compressed latent representations, where high-frequency garment details are inevitably lost during VAE encoding. Consequently,(b) latentspace generation struggles to faithfully preserve fine-grained garment textures. (c) Our framework leverages the VAR model for efficient holistic structure synthesis and dense image matching for pixel-space garment detail recovery, enabling both fast inference and faithful detail preservation.  \nshape, ii) well preserve person attributes including appearance and pose, and iii) faithfully retain garment details such as logos and text. Early GAN-based methods (Han et al., 2018; Wang et al., 2018; Choi et al., 2021) typically follow a two-stage pipeli","cbCaiavQ10qfrEWb","https://ap.wps.com/l/cbCaiavQ10qfrEWb","pdf",27294026,3,1,15,"English","en",105,"# Abstract\n# Introduction","[{\"question\":\"What problem does the document address in virtual try-on (VTON)?\",\"answer\":\"It targets efficient and high-fidelity VTON that preserves the person while accurately deforming the garment and synthesizing fine garment details.\"},{\"question\":\"Why do diffusion-based VTON methods struggle with high detail fidelity?\",\"answer\":\"They incur high inference latency and lose high-frequency garment details due to inherent compression in the VAE latent space.\"},{\"question\":\"What is the key idea behind STAR-VTON compared with earlier approaches?\",\"answer\":\"STAR-VTON decouples latent-space structural synthesis from pixel-space fine-detail recovery, using a matching-informed refiner to map dense garment details from the source.\"}]",1784209243,38,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"structure-detail-decoupled-autoregressive-generation-for-fast-and-high-fidelity-virtual-try-on","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/structure-detail-decoupled-autoregressive-generation-for-fast-and-high-fidelity-virtual-try-on/86188/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the document address in virtual try-on (VTON)?","Question",{"text":75,"@type":76},"It targets efficient and high-fidelity VTON that preserves the person while accurately deforming the garment and synthesizing fine garment details.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Why do diffusion-based VTON methods struggle with high detail fidelity?",{"text":80,"@type":76},"They incur high inference latency and lose high-frequency garment details due to inherent compression in the VAE latent space.",{"name":82,"@type":73,"acceptedAnswer":83},"What is the key idea behind STAR-VTON compared with earlier approaches?",{"text":84,"@type":76},"STAR-VTON decouples latent-space structural synthesis from pixel-space fine-detail recovery, using a matching-informed refiner to map dense garment details from the source.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]