[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83395-en":3,"doc-seo-83395-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83395,13056703020460,"Valentina","https://ap-avatar.wpscdn.com/avatar/be000253dac470eee5d?_k=1778207105932848923",8,"Research & Report","Structural Bottlenecks on Frequency Representation in End-to-End Audio Models","End-to-end neural audio models can compress and generate high-fidelity audio, yet strong results do not guarantee interpretable internal variables like pitch or timbre. The study analyzes whether modern strided convolutional encoders preserve access to frequency-localized primitives that underlie such features. It identifies two architecture- and signal-structured bottlenecks: alias-induced collapse that limits representational capacity and filter separability loss that constrains frequency resolution. Controlled experiments measure 31–35% collapse rates and 9–35× excess filter bandwidth, and introduce Gabor Latent Refactorization (GLRF) to restore frequency-localized access.","arXiv :2607 .08545v 1 [ cs . SD] 9 Jul 2026  \nStructural Bottlenecks on Frequency Representation in End-to-End Audio Models  \nNicole Cosme-Clifford  \nYale University  \nNew Haven, CT  \n[nicole.cosme@yale.edu](nicole.cosme@yale.edu)  \nAbstract  \nEnd-to-end neural audio models achieve high-fidelity compression and generation.  \nWe might read that performance as evidence they directly represent interpretable features such as pitch and timbre, but a model can produce plausible outputs without doing so. A model may encode these features in any reachable basis, but regardless of which, the features are well described as compositions of timefrequency-localized primitives. Whether state-of-the-art encoders preserve access to these primitives, and thus to compositions of them, remains unclear. Through theoretical analysis and controlled experiments, we show that several state-ofthe-art strided convolutional encoders impose two structural bottlenecks, both predictable from architecture and signal structure, on access to these primitives:  \n(1) they collapse primitives into alias equivalence classes, establishing a bound on representational capacity, and (2) they limit the frequency resolution available to learned filters, restricting separability. For well structured data, we find collapse rates of 31-35% and filter bandwidths 9-35x above the theoretical resolution bound, confirming that both bottlenecks arise under realistic signal conditions. We then introduce Gabor Latent Refactorization (GLRF), a lightweight post-hoc intervention that re-expresses encoder latents in a frequency-localized basis, reducing filter bandwidths from 10–35x to 1.5–3x of the theoretical resolution bound while preserving reconstruction fidelity and improving control over attributes like pitch.  \nThese results show that the encoders in question predictably degrade access to frequency-localized primitives, entangling the features that depend on them, and that a lightweight, retraining-free intervention can recover much of that access, improving steerability and interpretability.  \n1 Introduction  \nModern neural audio systems achieve state-of-the-art performance across compression and generation, with scaling consistently improving empirical results Défossez et al. [2022], Kumar et al. [2023], Evans et al. [2024, 2025] . It is tempting to read that strong performance as evidence that interpretable features like pitch and timbre are directly encoded in these systems’ learned representations. Yet, interpretability studies have largely not recovered them Singh et al. [2025], Beguš and Zhou [2022], Lee et al. [2017], Vu et al. [2024], Muckenhirn et al. [2019] . Have these models learned interpretable time-frequency structure, or can strong performance exist without it? We argue the latter: state-ofthe-art encoders foreclose access to the primitives that underlie these features, and the features do notreappear in some alternative learned form.  \nMany physical sound-generating systems are, over short timescales, well described by normal modes Oppenheim et al. [1999], Morse and Ingard [1986] . Sound can then be locally approximated as a superposition x(t) = Pi ci (t), where each ci (t) is a narrowband oscillation centered at a distinct frequency, fi. Human auditory perception exploits this structure, as human-meaningful features (e.g.,  \nPreprint.  \npitch, timbre) follow from relationships between these time-frequency-localized primitives Bregman [1994] . Whether artificial systems preserve access to these primitives (and compositions of them) is determined in part by encoder architecture.  \nState-of-the-art models increasingly operate end-to-end over raw waveforms, using strided convolutional encoders to compress high-dimensional signals into compact latents Défossez et al. [2022], Kumar et al. [2023], Evans et al. [2024, 2025], Zeghidour et al. [2021a], Dhariwal et al. [2020] . We show that this class of encoders exhibits two structural bottlenecks on structured contr","cbCaiqzRSTqTz0dk","https://ap.wps.com/l/cbCaiqzRSTqTz0dk","pdf",505389,2,1,20,"English","en",105,"# Introduction\n## Frequency-localized primitives and interpretability\n## Structural bottlenecks from strided convolutional encoders\n## Practical stakes for controllable audio generation\n## Proposed intervention: GLRF","[{\"question\":\"What question does the paper address about end-to-end audio models?\",\"answer\":\"Whether high performance implies that interpretable features such as pitch and timbre are directly represented, or whether strong output quality can occur even when the underlying frequency-localized structure is not accessible.\"},{\"question\":\"What are the two structural bottlenecks identified in the strided convolutional encoders?\",\"answer\":\"First, alias-induced collapse that groups distinct frequency components into indistinguishable proxy components, limiting representational capacity. Second, separability failure where learned filters operate with frequency resolution well above the theoretical limit set by the encoder’s receptive field.\"},{\"question\":\"How does Gabor Latent Refactorization (GLRF) help?\",\"answer\":\"GLRF is a lightweight post-hoc intervention that re-expresses encoder latents in a frequency-localized basis, reducing filter bandwidths from roughly 10–35× to about 1.5–3× of the theoretical resolution bound while preserving reconstruction fidelity and improving control over attributes like pitch.\"}]",1784187210,50,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"structural-bottlenecks-on-frequency-representation-in-end-to-end-audio-models","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/structural-bottlenecks-on-frequency-representation-in-end-to-end-audio-models/83395/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-22","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What question does the paper address about end-to-end audio models?","Question",{"text":75,"@type":76},"Whether high performance implies that interpretable features such as pitch and timbre are directly represented, or whether strong output quality can occur even when the underlying frequency-localized structure is not accessible.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What are the two structural bottlenecks identified in the strided convolutional encoders?",{"text":80,"@type":76},"First, alias-induced collapse that groups distinct frequency components into indistinguishable proxy components, limiting representational capacity. Second, separability failure where learned filters operate with frequency resolution well above the theoretical limit set by the encoder’s receptive field.",{"name":82,"@type":73,"acceptedAnswer":83},"How does Gabor Latent Refactorization (GLRF) help?",{"text":84,"@type":76},"GLRF is a lightweight post-hoc intervention that re-expresses encoder latents in a frequency-localized basis, reducing filter bandwidths from roughly 10–35× to about 1.5–3× of the theoretical resolution bound while preserving reconstruction fidelity and improving control over attributes like pitch.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,114,119,122,126,129,133],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":29,"slug":113},6,"Technology","technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":22,"slug":125},9,"Religion & Spirituality","religion-spirituality",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":127,"show_sort_weight":22,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":46,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":46,"category_name":135,"show_sort_weight":106,"slug":136},19,"General","general"]