[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82784-en":3,"doc-seo-82784-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},82784,137441390410,"Hazel","https://ap-avatar.wpscdn.com/avatar/2000252f4ab5702993?_k=1776741390130283984",8,"Research & Report","Speaker-Disentangled Chunk-Wise Regression for Syllabic Tokenization","Unsupervised syllabic tokenization seeks discrete syllabic units that reflect latent linguistic structure from raw speech. Existing approaches distill pretrained HuBERT representations into syllabic segments, but an utterance-level cross-entropy objective can drive the model to predict speaker identity instead of linguistic content, reducing token purity. A speaker-disentangled syllabic tokenizer is proposed, regressing speaker-perturbed student representations to clean teacher targets within fixed-length chunks. Experiments show state-of-the-art results for syllable boundary detection and segment clustering, and improved LM understanding over phone-level baselines.","Received XX Month, XXXX; revised XX Month, XXXX; accepted XX Month, XXXX; Date of publication XX Month, XXXX; date of  \ncurrent version XX Month, XXXX.  \nDigital Object Identifier 10.1109/XXXX.2022.1234567  \nSpeaker-Disentangled Chunk-Wise Regression for Syllabic Tokenization  \narXiv :2607 .04064v 1 [ cs .CL] 5 Jul 2026  \nRYOTA KOMATSU1 , KOTA KAWAKITA1 , TAKUMA OKAMOTO2 (Member, IEEE), AND TAKAHIRO SHINOZAKI1 (Member, IEEE)  \n1 Institute of Science Tokyo, Meguro, Tokyo 152-8550, Japan  \n2 National Institute of Information and Communications Technology, Kyoto 619-0289, Japan Corresponding author: Ryota Komatsu (email: [komatsu.r.ab@m.titech.ac.jp](komatsu.r.ab@m.titech.ac.jp)).  \nThis work was supported in part by JTEKT Corporation and in part by JSPS KAKENHI under Grant JP22K12069 .  \nABSTRACT Unsupervised syllabic tokenization aims to learn discrete syllabic tokens that capture latent linguistic content-related structure from raw speech. Recent syllabic tokenization methods employ teacherstudent distillation of the pretrained HuBERT to organize latent speech frame representations into syllabic segments. However, when trained with an utterance-level cross-entropy objective, the model predicts speaker identity rather than linguistic content, thereby compromising the purity of syllabic tokens. To address this problem, we propose a speaker-disentangled syllabic tokenizer that regresses speaker-perturbed student representations toward clean teacher targets within fixed-length chunks. Experimental results demonstrate that our proposed method achieves state-of-the-art performance in syllable boundary detection and syllabic segment clustering. Moreover, a speech language model trained on our syllabic tokens achievesa 7% relative improvement in syntactic and semantic understanding over the phone-level SpiRit-LM.  \nINDEX TERMS Self-supervised learning, speech language models, speech tokenization, syllable discovery.  \nI. INTRODUCTION  \nSELF-SUPERVISED speech representation learning has  \nbeen shown to effectively extract phonetic content from raw speech [1], [2] . This enables phonetic tokens to serve as pseudo-transcripts, thereby allowing language modeling directly on speech tokens [3] . As a result, speech language models (LMs) offer a unified framework for understanding and generating spoken language, and have emerged as a foundation for spoken dialogue modeling [4]–[7] .  \nTo transfer linguistic knowledge from text LMs to speech LMs, SpiRit-LM introduces word-level speech-text interleaving, where textually pretrained LMs are continually trained on sequences that alternate between phonetic and text tokensat word boundaries [8] . However, a fundamental challenge lies in the mismatch of token granularity between speech and text. Learned phonetic tokens typically occur at a high frame rate (12.5–50 Hz), whereas text is encoded using coarser subword tokens. This lower linguistic information density in speech tokens reduces computational efficiency and exacerbates the granularity mismatch, which can hinder speech-text alignment.  \nTo mitigate this mismatch, recent approaches aim to discover linguistically meaningful, coarser syllabic tokens [9]–[13] . In particular, Cho et al. proposed SD-HuBERT, a selfdistillation framework for the pretrained HuBERT based on an utterance-level cross-entropy objective [10],[14]. This approach implicitly organizes latent frame representations into syllabic segments in an intermediate Transformer layer [15], as shown in Figure 1b [10] . Syllabic tokens are derived as cluster indices via a three-step tokenization procedure: 1) computing a self-similarity matrix of frame-level features, 2) segmenting the matrix to identify syllable boundaries, and 3) quantizing the segment-wise average features. Recent studies have shown that speech LMs built on these syllabic tokens outperform speech LMs trained on phone-level tokens in syntactic understanding, suggesting that they better capture linguistically abstract c","cbCaivD0aUFXvC1o","https://ap.wps.com/l/cbCaivD0aUFXvC1o","pdf",1018402,3,1,10,"English","en",105,"# Introduction\n## Motivation: Speech-text token granularity mismatch\n## Prior work: SD-HuBERT and syllable discovery\n## Observed issues: prototype collapse and speaker dominance\n## Proposed direction: speaker-disentangled frame-wise regression","[{\"question\":\"What problem does the paper target in unsupervised syllabic tokenization?\",\"answer\":\"It targets the tendency of existing syllabic tokenization models to learn tokens contaminated by speaker identity rather than linguistic content, reducing the purity of syllabic tokens.\"},{\"question\":\"How does the proposed method differ from utterance-level cross-entropy distillation?\",\"answer\":\"It regresses speaker-perturbed student representations toward clean teacher targets using a chunk-wise, fixed-length regression objective rather than relying on utterance-level classification.\"},{\"question\":\"What evidence is reported to validate the approach?\",\"answer\":\"Experimental results show state-of-the-art performance in syllable boundary detection and syllabic segment clustering, and a speech language model trained on the syllabic tokens yields a 7% relative improvement in syntactic and semantic understanding over a phone-level baseline.\"}]",1784182913,25,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"speaker-disentangled-chunk-wise-regression-for-syllabic-tokenization","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/speaker-disentangled-chunk-wise-regression-for-syllabic-tokenization/82784/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-23","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the paper target in unsupervised syllabic tokenization?","Question",{"text":75,"@type":76},"It targets the tendency of existing syllabic tokenization models to learn tokens contaminated by speaker identity rather than linguistic content, reducing the purity of syllabic tokens.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the proposed method differ from utterance-level cross-entropy distillation?",{"text":80,"@type":76},"It regresses speaker-perturbed student representations toward clean teacher targets using a chunk-wise, fixed-length regression objective rather than relying on utterance-level classification.",{"name":82,"@type":73,"acceptedAnswer":83},"What evidence is reported to validate the approach?",{"text":84,"@type":76},"Experimental results show state-of-the-art performance in syllable boundary detection and syllabic segment clustering, and a speech language model trained on the syllabic tokens yields a 7% relative improvement in syntactic and semantic understanding over a phone-level baseline.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":22,"slug":133},"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]