[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85907-en":3,"doc-seo-85907-105":29,"detail-sidebar-cat-0-en-105":90},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":11,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},85907,7971461740886,"Theodore","https://ap-avatar.wpscdn.com/davatar_3d24733baf745e90a7e4bdd5f77d97b2",8,"Research & Report","FdAudio MeanFlow-Anchored Frchet-Distance Post-Training for One-Step Text-to-Audio Generation","One-step text-to-audio generation often delivers faster inference but falls short of multi-step diffusion and flow-matching models in perceptual and distributional quality. FdAudio bridges this gap with a post-training method that optimizes the final one-step output distribution via multirepresentation Frchet-distance (FD) loss across pretrained embedding spaces. To avoid multi-step degradation caused by naive FD post-training, FdAudio introduces a MeanFlow consistency anchor that preserves the denoising trajectory. Experiments show state-of-the-art few-step one-step quality and improved FD/FAD metrics while maintaining high-fidelity audio for longer sampling paths with reduced latency.","FdAudio: MeanFlow-Anchored Frchet-Distance Post-Training for One-Step Text-to-Audio  \nGeneration  \nKuan-Po Huang⋆ , Bo-Ru Lu†, Ho-Lam Chung⋆ , Shih-Hsin Wang⋆ , Hung-yi Lee⋆  \n⋆ National Taiwan University †Amazon  \narXiv :2607 . 10421v1 [ ee ss .AS] 11 Jul 2026  \nAbstract—While recent few-step sampling text-to-audio generation models like MeanAudio substantially accelerate generation by modeling average velocities, their strict one-step generation quality still lags significantly behind multi-step counterparts. We propose FdAudio to bridge this gap. Unlike MeanAudio, which relies solely on regression against target velocity fields, our post-training approach optimizes the final one-step distribution directly across pre-trained embedding spaces via a multirepresentation Frchet-distance (FD) loss. Crucially, to prevent the multi-step degradation that naive post-training with FDloss causes, we introduce a MeanFlow consistency objective as a structural anchor. Results demonstrate that FdAudio establishes state-of-the-art one-step T2A generation quality among few-step systems, yielding an 11.4% reduction in FD score and a 28.8% improvement in FAD score relative to the baseline MeanAudio framework. Notably, we solve FD post-training’s naive multi-step degradation issue by proposing the MeanFlow anchor, enabling a 25-step sampling path to maintain high-fidelity audio synthesis that matches or surpasses strong multi-step models at a fraction of their computational latency.  \nIndex Terms—text-to-audio generation, one-step generation, MeanFlow, Frchet-distance, flow-matching  \nI. INTRODUCTION  \nText-to-audio (T2A) aims to automate the process by synthesizing general sound, from ambient scenes to sound events, directly from a natural language prompt, with applications across games, film, and digital content creation. Because inference is performed far more frequently than training in these practical applications, the inference cost becomes a central concern [1] . However, state-of-the-art systems are mostly built on diffusion [2] and flow-matching [3] models [4]–[11], which produce high-fidelity audio through an iterative sampling process. The model is evaluated sequentially over tens to hundreds of timesteps, at each time step taking a partially denoised latent and predicting a small update that gradually recovers the final clean audio latents. This sequential loop is the primary source of their high inference latency.  \nTo remove this bottleneck, recent systems distill or shorten the sampling trajectory for few-or one-step inference, including ConsistencyTTA [12], AudioLCM [13], AudioTurbo [14], SoundCTM [15], AudioDEAR [16], and MeanAudio [17] . Although dramatically faster, these few-step models still trail their multi-step counterparts in generation quality. This motivates the central question of our work: can we substantially improve the generation quality of one-step audio generation?  \n†This work is unrelated to the author’s position at Amazon.  \nA recent promising direction in the image generation domain is to apply Frchet-distance (FD) post-training [18] to finetune a pretrained one-step generator [19], [20] so that the distribution of its one-step outputs matches real data in the feature spaces of pretrained representation extractors [21]–[26] . This approach is appealing because it is simple and direct: it only requires precomputed real-data statistics, without teacher distillation, adversarial training, or per-sample regression targets.  \nWe adapt this recipe to audio by computing the FD-loss over pretrained audio encoders, including PANNs [27], PaSST [28], BEATs [29], and AudioMAE [30] . While this substantially improves one-step audio quality, we find that it severely harms multi-step sampling. For example, as shown in the results of Section V-B, a model post-trained with FD-loss achieves a strong one-step FAD of 1.27, but its FAD nearly triples to 3.61 when sampled with 25 steps. This suggests that FD post-training, when ","cbCaiiD6l27t1Cz4","https://ap.wps.com/l/cbCaiiD6l27t1Cz4","pdf",416430,4,1,"English","en",105,"# Introduction\n# Approach and Motivation\n# Frchet-distance Post-Training for One-Step Audio\n# MeanFlow Consistency Anchor\n# Contributions","[{\"question\":\"What problem does FdAudio address in one-step text-to-audio generation?\",\"answer\":\"FdAudio targets the quality gap between fast one-step generation and slower multi-step models, where one-step outputs lag in fidelity and distributional measures.\"},{\"question\":\"How does FdAudio use Frchet-distance (FD) in post-training?\",\"answer\":\"FdAudio performs post-training by optimizing the final one-step output distribution directly across pretrained representation embedding spaces using a multirepresentation FD loss.\"},{\"question\":\"Why does naive FD post-training harm multi-step sampling, and how is it fixed?\",\"answer\":\"Naive FD post-training can collapse the denoising trajectory into a direct noise-to-data mapping, making multi-step sampling less effective; MeanFlow consistency anchors the average velocity over sub-intervals to preserve the trajectory.\"}]",1784207081,20,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":85,"head_meta":87,"extra_data":89,"updated_unix":27},"fdaudio-meanflow-anchored-frchet-distance-post-training-for-one-step-text-to-audio-generation","",{"@graph":35,"@context":84},[36,52,67],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":20},"https://docshare.wps.com/document/fdaudio-meanflow-anchored-frchet-distance-post-training-for-one-step-text-to-audio-generation/85907/",{"url":51,"name":13,"@type":53,"author":54,"headline":13,"publisher":56,"fileFormat":59,"inLanguage":23,"description":14,"dateModified":60,"datePublished":61,"encodingFormat":59,"isAccessibleForFree":62,"interactionStatistic":63},"DigitalDocument",{"name":9,"@type":55},"Person",{"url":40,"name":57,"@type":58},"DocShare","Organization","application/pdf","2026-07-26","2026-07-16",true,{"@type":64,"interactionType":65,"userInteractionCount":20},"InteractionCounter",{"@type":66},"ViewAction",{"@type":68,"mainEntity":69},"FAQPage",[70,76,80],{"name":71,"@type":72,"acceptedAnswer":73},"What problem does FdAudio address in one-step text-to-audio generation?","Question",{"text":74,"@type":75},"FdAudio targets the quality gap between fast one-step generation and slower multi-step models, where one-step outputs lag in fidelity and distributional measures.","Answer",{"name":77,"@type":72,"acceptedAnswer":78},"How does FdAudio use Frchet-distance (FD) in post-training?",{"text":79,"@type":75},"FdAudio performs post-training by optimizing the final one-step output distribution directly across pretrained representation embedding spaces using a multirepresentation FD loss.",{"name":81,"@type":72,"acceptedAnswer":82},"Why does naive FD post-training harm multi-step sampling, and how is it fixed?",{"text":83,"@type":75},"Naive FD post-training can collapse the denoising trajectory into a direct noise-to-data mapping, making multi-step sampling less effective; MeanFlow consistency anchors the average velocity over sub-intervals to preserve the trajectory.","https://schema.org",{"og:url":51,"og:type":86,"og:title":13,"og:site_name":57,"og:description":14},"article",{"robots":88,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":91},[92,96,100,104,109,114,119,122,126,129,133],{"id":21,"doc_module":4,"doc_module_name":45,"category_name":93,"show_sort_weight":94,"slug":95},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":97,"show_sort_weight":98,"slug":99},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":45,"category_name":101,"show_sort_weight":102,"slug":103},"Exam",70,"exam",{"id":105,"doc_module":4,"doc_module_name":45,"category_name":106,"show_sort_weight":107,"slug":108},5,"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":45,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":28,"slug":125},9,"Religion & Spirituality","religion-spirituality",{"id":28,"doc_module":4,"doc_module_name":45,"category_name":127,"show_sort_weight":28,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":45,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":45,"category_name":135,"show_sort_weight":105,"slug":136},19,"General","general"]