[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82212-en":3,"doc-seo-82212-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},82212,1374391974468,"Eden","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","ReGen Hierarchical Multi Prompt Representation Generation for Efficient Waveform Diffusion Models","Representation alignment (REPA) is used to speed up diffusion training, but regularizing intermediate representations in diffusion Transformers can entangle latents and restrict generative capacity. ReGen introduces a hierarchical multi-prompt representation generation framework that jointly estimates multiple vector fields for both representations and data in one diffusion model. Generalized flow matching (GFM) improves conditional flow matching generalization. Experiments cover waveform diffusion models, neural audio codecs, and a LDM-based text-to-speech system, with efficient training and strong intelligibility and similarity.","ReGen: Hierarchical Multi-Prompt Representation Generation for Efficient Waveform Diffusion Models  \nSang-Hoon Lee 1 Ha-Yeong Choi 2  \narXiv :2607 .09 134v 1 [ cs . SD] 10 Jul 2026  \nAbstract  \nRepresentation alignment (REPA) has been investigated to accelerate diffusion training, but we observe that regularizing intermediate representations in diffusion Transformers (DiT) may implicitly entangle latents and limit generative capacity.  \nTo address this issue, we propose ReGen, a hierarchical multi-prompt representation generation framework that jointly estimates multiple vector fields for both representations and data within a single diffusion model. We further introduce generalized flow matching (GFM) to improve the generalization of conditional flow matching (CFM) .  \nWe validate ReGen on single-stage waveform diffusion models including neural audio codec and Wave-VAE. ReGen significantly improves waveform generation quality from highly compressed latent representations at 12.5 Hz. We also present ReGenVoice, a latent diffusion model (LDM) -based text-to-speech model that achieves strong speech intelligibility (WER) and speaker similarity (SIM) with a small dataset. Moreover, operating the LDM at 6.25 Hz with rich semantic and acoustic latent representation enables efficient training and sampling, requiring only 1 day of training on 4 GPUs and fast inference with an RTF of 0.08. Audio samples are available at [https:](https:)// [regenvoice.github.io/demo/](regenvoice.github.io/demo/) .  \n1. Introduction  \nRecently, audio has been getting interest in human-centric artificial intelligence systems, enabling both audio understanding and generation in spoken dialogue systems. To do this, it is essential to compress high-resolution raw waveform signals into low-resolution latent representations via vec-  \n1Department of Artificial Intelligence, Ajou University, Suwon, Korea 2 KT Corp., Seoul, Korea. Correspondence to: Sang-Hoon Lee \u003C[sanghoonlee@ajou.ac.kr](sanghoonlee@ajou.ac.kr) >.  \nProceedings of the 43 rd International Conference on Machine Learning, Seoul, South Korea. PMLR 306, 2026 . Copyright 2026 by the author(s) .  \ntor quantization (VQ) or variational autoencoders (VAEs), followed by large language models that predict VQ-based discrete audio tokens and diffusion models that generate VAE-based continuous audio vectors. Then, audio decoder converts these representation into an audible waveform.  \nWhile generative adversarial networks (GAN)-based neural audio codecs (Zeghidour et al., 2021) are commonly used, they have several drawbacks in extremely low-bitrate scenarios: 1) limited semantic capacity, which reduces speech intelligibility; 2) limited acoustic capacity, which degrades high-frequency details. Basically, semantic distillation with residual vector quantization (RVQ) (Zhang et al., 2024 ; Dfossez et al., 2024) has been widely used to improve the semantic consistency. However, it requires additional quantization and can reduce generative capacity.  \nMeanwhile, conditional flow matching (CFM)-based models have gained increasing attention as a new generation paradigm for audio modeling (Lee et al., 2025 ; Liu et al., 2025 ; Welker et al., 2025 ; Yao et al., 2025) . PeriodWaveTurbo (Lee et al., 2024) demonstrated that CFM-based pretraining and adversarial post-training significantly improve the performance and reduce overall training times. StreamFlow (Choi & Lee, 2025) introduces streaming flow matching for streaming DiT-based waveform generation, and analyzes representation alignment (REPA) (Yu et al., 2025) learning in diffusion Transformers (DiT) (Peebles & Xie, 2023) for waveform generation, where REPA accelerates training speed and can improve generative capacity via robust semantic alignment. However, recent studies (Wanget al., 2025d) report that REPA can suffer from capacity mismatch that decreases the capacity on generative ability due to implicitly entangled latent representations in DiT.  \nTo address th","cbCaiouVPvBYcVva","https://ap.wps.com/l/cbCaiouVPvBYcVva","pdf",552846,1,13,"English","en",105,"# Abstract\n# Introduction\n# Neural Waveform Generation with SSL\n# Contributions","[{\"question\":\"What problem does ReGen address in diffusion Transformers for waveform generation?\",\"answer\":\"ReGen targets the limitation of representation alignment (REPA), where regularizing intermediate representations can entangle latents and reduce generative capacity.\"},{\"question\":\"How does ReGen generate hierarchical multi-prompt representations?\",\"answer\":\"ReGen builds a hierarchical DiT that progresses from semantics to waveform and uses a multi-prompting mechanism guided through a masked-infilling strategy to jointly generate representations and waveform.\"},{\"question\":\"What is generalized flow matching (GFM) and how does it help training?\",\"answer\":\"GFM improves robustness of waveform-level flow matching by introducing a repulsive term in the vector field space, mitigating zero-collapse and enhancing generalization.\"}]",1784178853,33,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"regen-hierarchical-multi-prompt-representation-generation-for-efficient-waveform-diffusion-models","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/regen-hierarchical-multi-prompt-representation-generation-for-efficient-waveform-diffusion-models/82212/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does ReGen address in diffusion Transformers for waveform generation?","Question",{"text":75,"@type":76},"ReGen targets the limitation of representation alignment (REPA), where regularizing intermediate representations can entangle latents and reduce generative capacity.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does ReGen generate hierarchical multi-prompt representations?",{"text":80,"@type":76},"ReGen builds a hierarchical DiT that progresses from semantics to waveform and uses a multi-prompting mechanism guided through a masked-infilling strategy to jointly generate representations and waveform.",{"name":82,"@type":73,"acceptedAnswer":83},"What is generalized flow matching (GFM) and how does it help training?",{"text":84,"@type":76},"GFM improves robustness of waveform-level flow matching by introducing a repulsive term in the vector field space, mitigating zero-collapse and enhancing generalization.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]