[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85491-en":3,"doc-seo-85491-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},85491,962075006959,"Anda","https://ap-avatar.wpscdn.com/avatar/e0002397efbe92a78e?_k=1776741047341049297",8,"Research & Report","MixFlow Training: Alleviating Exposure Bias with Slowed Interpolation Mixture","MixFlow addresses the training–testing discrepancy, or exposure bias, in diffusion models by changing what the prediction network sees during training versus sampling. In standard training, each timestep input uses ground-truth noisy data as an interpolation of noise and clean data, while testing uses generated noisy data. MixFlow leverages the Slow Flow observation that the nearest ground-truth interpolation to generated noise corresponds to a higher-noise “slowed timestep.” It trains on a slowed interpolation mixture to reduce error accumulation and drift, improving class-conditional image and text-to-image generation, including strong ImageNet FID scores on RAE.","MixFlow Training: Alleviating Exposure Bias with Slowed Interpolation Mixture  \nHui Li 1 ,2 ,5 Fu-Yun Wang3 Haoyuan Xia 1 Jiayue Lyu2 Kaihui Cheng2  \nSiyu Zhu 1 ,2 ,5 B Jingdong Wang4 B  \n1 Shanghai Innovation Institute 2Fudan University 3The Chinese University of Hong Kong  \n4Baidu 5 Shanghai Academy of AI for Science  \n[https://mixflowgen.github.io/](https://mixflowgen.github.io/)  \narXiv :2512 . 19311v2 [ cs .CV] 12 Jul 2026  \nAbstract  \nThis paper studies the training-testing discrepancy (a.k.a. exposure bias) problem for improving the diffusion models. During training, the input of a prediction network at one training timestep is the corresponding ground-truth noisy data that is an interpolation of the noise and the data, and during testing, the input is the generated noisy data. We present a novel training approach, named MixFlow, for improving the performance. Our approach is motivated by the Slow Flow phenomenon: the ground-truth interpolation that is the nearest to the generated noisy data at a given sampling timestep is observed to correspond to a higher-noise timestep (termed slowed timestep), i.e., the corresponding ground-truth timestep is slower than the sampling timestep. MixFlow leverages the interpolations atthe slowed timesteps, named slowed interpolation mixture, for post-training the prediction network for each training timestep. Experiments over class-conditional image generation (including SiT, REPA, and RAE) and text-to-image generation validate the effectiveness of our approach. Our approach MixFlow over the RAE models achieve strong generation results on ImageNet: 1.43 FID (without guidance) and 1.10 (with guidance) at 256 × 256, and 1.55 FID (without guidance) and 1.10 (with guidance) at 512 × 512.  \n1. Introduction  \nWe study the training-testing discrepancy problem [38, 43], also known as exposure bias [15, 25, 27, 33, 34], for diffusion and flow matching models. During training, diffusion models learn a prediction network, where the input to the prediction network at each training timestep is the corresponding ground-truth noisy data, i.e., an interpolation of the noise and the data. During testing, the input to the prediction network is the generated noisy data. The difference of the inputs to the prediction network for training and testing,  \nB Corresponding authors  \n(a) (b)  \nFigure 1 . Illustrating (1) the Slow Flow phenomenon during the sampling process: the timestep (y-axis), corresponding to the ground truth noisy data that is the nearest to the generated noisy data at the sampling timestep t (x-axis), is slower (with higher noise), i.e., the shading area is under the line x = y; and (2) the effectiveness of MixFlow training: the range of slowed timesteps for (b) MixFlow training is smaller and closer to the sampling steps than (a) standard training, indicating that MixFlow training effectively alleviates the training-testing discrepancy. The boundary of the shading area in (b) is plotted as blue lines in (a) . Note: x-axisthe sampling timestep at which the noisy data is generated; y-axisthe slowed timestep corresponding to the ground truth noisy data that is the nearest to the generated noisy data; shading area-the range (the vertical line) of slowed timesteps at each sampling step; noise corresponds to timestep 0, and data corresponds to timestep 1. The slowed timestep ranges are obtained from 20, 000 training images in ImageNet [1], 50 sampling steps, and SiT-B [31] . Details on how to plot the figures are provided in Appendix A.  \ni.e., the training-testing discrepancy, is one of the reasons leading to the prediction discrepancy and accordingly the problems of error accumulation and sampling drift.  \nThere are two main lines of solutions to alleviating the discrepancy problem. One line is to modify the training procedure [15, 33] . For example, Input Perturbation [33] conducts an input perturbation on the ground truth noisy data, and self-forcing [15] uses the generated noisy data asthe","cbCaidY8teFx6gKH","https://ap.wps.com/l/cbCaidY8teFx6gKH","pdf",27261925,1,23,"English","en",105,"# Introduction\n## Training-testing discrepancy (exposure bias)\n## Existing solution lines: training vs sampling modifications\n## Slow Flow motivation and MixFlow overview","[{\"question\":\"What problem does MixFlow target in diffusion models?\",\"answer\":\"MixFlow targets the training–testing discrepancy, commonly called exposure bias, which causes error accumulation and sampling drift due to mismatched inputs between training and testing.\"},{\"question\":\"What is the Slow Flow phenomenon used to motivate MixFlow?\",\"answer\":\"During sampling, the ground-truth noisy data nearest to the generated noisy data at timestep t corresponds to a higher-noise “slowed timestep” mt, meaning the associated ground-truth timestep is slower than t.\"},{\"question\":\"How does MixFlow change the training procedure?\",\"answer\":\"MixFlow trains the prediction network using ground-truth slowed interpolations: each training timestep input becomes a mixture of interpolations from slowed timesteps, formed by selecting slowed interpolations and sampling the training timestep with a Beta(2,1) distribution.\"}]",1784203988,58,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"mixflow-training-alleviating-exposure-bias-with-slowed-interpolation-mixture","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/mixflow-training-alleviating-exposure-bias-with-slowed-interpolation-mixture/85491/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does MixFlow target in diffusion models?","Question",{"text":75,"@type":76},"MixFlow targets the training–testing discrepancy, commonly called exposure bias, which causes error accumulation and sampling drift due to mismatched inputs between training and testing.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What is the Slow Flow phenomenon used to motivate MixFlow?",{"text":80,"@type":76},"During sampling, the ground-truth noisy data nearest to the generated noisy data at timestep t corresponds to a higher-noise “slowed timestep” mt, meaning the associated ground-truth timestep is slower than t.",{"name":82,"@type":73,"acceptedAnswer":83},"How does MixFlow change the training procedure?",{"text":84,"@type":76},"MixFlow trains the prediction network using ground-truth slowed interpolations: each training timestep input becomes a mixture of interpolations from slowed timesteps, formed by selecting slowed interpolations and sampling the training timestep with a Beta(2,1) distribution.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]