[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82920-en":3,"doc-seo-82920-105":30,"detail-sidebar-cat-0-en-105":83},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},82920,8796095461610,"Oliver","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","Towards Streaming Neural Speech Codecs through Time-Invariant Representations","Neural speech codecs are increasingly used as intermediate representations for codec-based speech generation. TiCodec proposes a factorized representation with a Time-Invariant Representation Extraction (TIRE) module, separating time-varying speech content from time-invariant information to reduce frame-level modeling needs. The study probes what TIRE captures and assesses its suitability for low-latency processing. Results from probing tasks and segment selection show improved robustness and reconstruction quality, culminating in streaming inference over 660ms blocks with minimal degradation.","arXiv :2607 .05250v 1 [ cs .CL] 6 Jul 2026  \nTowards Streaming Neural Speech Codecs through Time-Invariant Representations  \nK´elian Est`eve 1 , Salima Mhdaffar 1[0000−0002−8472−6890], Mickael Rouvier 1[0000−0003−3541−3385], Richard Dufour2[0000−0003−1203−9108], and Yannick Est`eve 1[0000−0002−3656−8883]  \n1 LIA, Avignon Universit´e, France [first.last@univ-avignon.fr](first.last@univ-avignon.fr)  \n2 LS2N, Nantes Universit´e, France [first.last@univ-nantes.fr](first.last@univ-nantes.fr)  \nAbstract. Neural speech codecs are increasingly used as intermediate representations in codec-based speech generation systems. TiCodec introduces a factorized representation that separates time-varying speech content from time-invariant information through a Time-Invariant Representation Extraction (TIRE) module, potentially reducing the amount of information that must be modeled at the frame-level.  \nIn this work, we investigate the nature of the information captured by TIRE representations and their suitability for low-latency speech processing. Using a series of probing tasks, we analyze the influence of the encoder layer and show that intermediate layers capture complementary speaker-and environment-related information while containing little linguistic content. We further study several segment selection strategies for TIRE training and demonstrate that cross-file sampling improves the robustness of invariant representations. Based on these findings, we propose Dual-TIRE, a multi-level architecture that exploits the complementarity of different encoder layers and improves speech reconstruction quality and speaker similarity.  \nFinally, we evaluate TiCodec in a streaming inference setting using successive 660ms processing blocks. Results show that streaming operation can be achieved without significant degradation in reconstruction performance, highlighting the potential of factorized neural codec representations for future low-latency speech generation systems.  \n1 Introduction  \nNeural speech codecs have become a key component of modern speech processing systems. Originally developed for low-bitrate speech compression, neural codecs such as SoundStream [1], EnCodec [2], and Descript Audio Codec (DAC) [3] learn compact discrete representations that preserve perceptual quality while significantly reducing the bitrate of speech signals. Beyond compression, these discrete speech units have recently emerged as a versatile representation for speech generation and multimodal language modeling.  \nThe availability of discrete speech representations has enabled a new generation of codec-based speech generation systems. Models such as AudioLM [4],  \n2 K. Est`eve et al.  \nSPEAR-TTS [5], VALL-E [6], and Voicebox [7] use neural codec tokens as intermediate representations for speech synthesis and spoken language generation. By leveraging techniques initially developed for large language models, these approaches have demonstrated remarkable capabilities in speech generation, voice cloning, and spoken dialogue modeling.  \nDespite these advances, a major challenge remains. Neural speech codecs typically produce large numbers of frame-level tokens, often distributed across multiple Residual Vector Quantization (RVQ) layers. Consequently, speech generation systems must predict hundreds of discrete units per second, resulting in long sequences, increased computational cost, and higher inference latency. This limitation is particularly problematic for real-time speech applications, where low-latency generation is essential.  \n1.1 Motivation for Streaming Speech Generation  \nThe challenge of efficient speech generation becomes even more critical in streaming scenarios such as spoken dialogue systems, conversational agents, simultaneous speech translation, and real-time text-to-speech synthesis. In these applications, speech must be generated continuously while only a limited amount of future context is available. The latency introduced by the prediction of la","cbCaitc6WdZ3bCst","https://ap.wps.com/l/cbCaitc6WdZ3bCst","pdf",465902,2,1,15,"English","en",105,"# Abstract\n# Introduction\n## Motivation for Streaming Speech Generation\n## Time-Invariant Representations for Neural Speech Coding","[{\"question\":\"How is streaming performance evaluated in the document?\",\"answer\":\"TiCodec is evaluated in a streaming inference setting using successive 660ms processing blocks, and streaming can be achieved with no significant degradation in reconstruction performance.\"}]",1784183963,38,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":78,"head_meta":80,"extra_data":82,"updated_unix":28},"towards-streaming-neural-speech-codecs-through-time-invariant-representations","",{"@graph":36,"@context":77},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/towards-streaming-neural-speech-codecs-through-time-invariant-representations/82920/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-23","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71],{"name":72,"@type":73,"acceptedAnswer":74},"How is streaming performance evaluated in the document?","Question",{"text":75,"@type":76},"TiCodec is evaluated in a streaming inference setting using successive 660ms processing blocks, and streaming can be achieved with no significant degradation in reconstruction performance.","Answer","https://schema.org",{"og:url":51,"og:type":79,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":81,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":84},[85,89,93,97,102,107,112,115,120,123,127],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":86,"show_sort_weight":87,"slug":88},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":90,"show_sort_weight":91,"slug":92},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Exam",70,"exam",{"id":98,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},5,"Comic",60,"comic",{"id":103,"doc_module":4,"doc_module_name":46,"category_name":104,"show_sort_weight":105,"slug":106},6,"Technology",50,"technology",{"id":108,"doc_module":4,"doc_module_name":46,"category_name":109,"show_sort_weight":110,"slug":111},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":113,"slug":114},30,"research-report",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},9,"Religion & Spirituality",20,"religion-spirituality",{"id":118,"doc_module":4,"doc_module_name":46,"category_name":121,"show_sort_weight":118,"slug":122},"World Cup","world-cup",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":124,"slug":126},10,"Lifestyle","lifestyle",{"id":128,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":98,"slug":130},19,"General","general"]