[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82149-en":3,"doc-seo-82149-105":28,"detail-sidebar-cat-0-en-105":90},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":11,"language":21,"language_code":22,"site_id":23,"html_lang":22,"table_of_contents":24,"faqs":25,"seo_title":13,"seo_description":14,"update_tm":26,"read_time":27},82149,1099514067438,"River Wang","https://ap-avatar.wpscdn.com/avatar/100002539ee87300030?x-image-process=image/resize,m_fixed,w_180,h_180&k=1780474512215547542",8,"Research & Report","Phone Segmentation and Recognition through Phonological Activation Mapping","Phone segmentation and recognition are closely linked tasks, yet many modern systems treat them separately. The approach argues that phonetic structure is already present in the latent representations of self-supervised speech models (S3Ms) and can be guided to solve both problems jointly. It uses S3M-based Phonological Activation Mapping (SPAM) to convert each frame embedding into phonological feature activations such as voicing and nasality. Lightweight prediction heads then perform recognition and boundary segmentation with under one minute of labeled transcriptions, generalizing to unseen phones and improving results across varied datasets.","Phone Segmentation and Recognition through Phonological Activation Mapping  \nShikhar Bharadwaj§∗ , Kwanghee Choi†∗ , Stephen McIntosh‡∗ , Chin-Jou Li§ , Eunjung Yeo†, Daisuke Saito‡, Nobuaki Minematsu‡, Shinji Watanabe§ , Jian Zhu¶ , David Harwath† and David R. Mortensen§  \n§ CMU, USA †UT Austin, USA ‡UTokyo, Japan ¶ UBC, Canada  \n{sbharad2,[dmortens](dmortens}@andrew.cmu.edu {kwanghee)[}](dmortens}@andrew.cmu.edu {kwanghee)[@andrew.cmu.edu](dmortens}@andrew.cmu.edu {kwanghee)[ {](dmortens}@andrew.cmu.edu {kwanghee)[kwanghee](dmortens}@andrew.cmu.edu {kwanghee),[harwath](harwath}@utexas.edu {smcintosh)[}](harwath}@utexas.edu {smcintosh)[@utexas.edu](harwath}@utexas.edu {smcintosh)[ {](harwath}@utexas.edu {smcintosh)[smcintosh](harwath}@utexas.edu {smcintosh),[mine](mine}@gavo.t.u-tokyo.ac.jp)[}](mine}@gavo.t.u-tokyo.ac.jp)[@gavo.t.u-tokyo.ac.jp](mine}@gavo.t.u-tokyo.ac.jp)  \narXiv :2607 .09020v1 [ ee ss .AS] 10 Jul 2026  \nAbstract—Phone segmentation and recognition are inherently related tasks, yet modern approaches typically model them separately. We argue that phonetic structure is already latent in the representations of self-supervised speech models (S3Ms), and one only needs to steer them to solve both tasks. We leverage S3Mbased Phonological Activation Mapping (SPAM), which maps each S3M representation frame to a vector of phonological feature activations, such as voicing and nasality. On top of SPAM, we introduce two simple but effective lightweight, gradient-descentfree prediction heads: a recognition head and a segmentation head. Our method requires less than a minute of phonetic transcriptions, and generalizes to unseen phones during training. Across a diverse range of datasets, our approach attains strong segmentation and recognition performance.  \nIndex Terms—phone segmentation, phone recognition, selfsupervised learning, phonetics, phonology  \nI. INTRODUCTION  \nPhones are the smallest independent units of speech sound, conventionally transcribed across languages with the International Phonetic Alphabet (IPA) . Locating and recognizing them, i.e., phone segmentation and recognition, yields timealigned phonetic transcriptions [1], [2] . These are foundationalto a range of applications, including clinical assessment of pathological speech, computer-assisted pronunciation training, and the documentation of endangered languages [3]–[8] .  \nHowever, obtaining such annotations is challenging, since expert phonetic transcription is slow and costly. Annotating one hour of speech can take a trained phonetician roughly 40 to 100 hours [9], [10] . Transcriptions can also be subjective, with even experienced transcribers sometimes disagreeing on phone labels and segment boundaries [11], [12] .  \nThese challenges have motivated automatic phone recognition and segmentation models. Modern phone recognition is typically framed as a sequence-to-sequence task, much like automatic speech recognition (ASR) [13]–[16] . Accordingly, it is often trained with a connectionist temporal classification (CTC) loss [13], [17] or an attention-based encoder-decoder architecture [15], [18], [19] . Phone segmentation, by contrast, is typically cast as frame-wise classification [20]–[23], labeling each frame as a boundary (1) or not (0) . Yet for human listeners, the two are inseparable: hearing speech, we perceive both which phones are spoken and where they occur in time. This  \n∗ These authors contributed equally. Author order determined by random shuffling.  \nFig. 1. Overview of S3M-based Phonological Activation Mapping (SPAM). Left: Phonological vectors [27], [28] can be found in S3M representation space by taking the difference of means. For example, vvoi+ is the difference between the mean representations of voiced phones and other phones.  \nRight: S3M frame representations rt and r t′ from different timesteps t and t′ are projected onto phonological vectors [8], [28] (e.g., vback+), yielding activation values (top right) . Stacking these proj","cbCaijG58lFI5xXC","https://ap.wps.com/l/cbCaijG58lFI5xXC","pdf",703824,1,"English","en",105,"# Introduction\n## Unified representation for segmentation and recognition\n## Phonological feature decomposition and SPAM\n## Modeling phonological vectors in S3M space\n# Method Overview\n## Segmentation head and phonological dissimilarity\n## Recognition head and joint prediction","[{\"question\":\"Why is phone segmentation and recognition often modeled separately in modern systems?\",\"answer\":\"Many approaches frame recognition and segmentation as different task types, such as sequence-to-sequence for recognition and frame-wise boundary classification for segmentation.\"},{\"question\":\"What is SPAM and what does it produce from S3M representations?\",\"answer\":\"SPAM maps each S3M representation frame to activations over phonological feature vectors (e.g., voicing, nasality), creating a time-aligned phonological activation map across frames.\"},{\"question\":\"How does the method achieve segmentation and recognition with limited supervision?\",\"answer\":\"It introduces lightweight recognition and segmentation heads on top of SPAM, requiring less than a minute of phonetic transcriptions and generalizing to unseen phones during training.\"}]",1784178452,20,{"code":4,"msg":29,"data":30},"ok",{"site_id":23,"language":22,"slug":31,"title":13,"keywords":32,"description":14,"schema_data":33,"social_meta":85,"head_meta":87,"extra_data":89,"updated_unix":26},"phone-segmentation-and-recognition-through-phonological-activation-mapping","",{"@graph":34,"@context":84},[35,52,67],{"@type":36,"itemListElement":37},"BreadcrumbList",[38,42,46,49],{"item":39,"name":40,"@type":41,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":43,"name":44,"@type":41,"position":45},"https://docshare.wps.com/document/","Document",2,{"item":47,"name":12,"@type":41,"position":48},"https://docshare.wps.com/document/research-report/",3,{"item":50,"name":13,"@type":41,"position":51},"https://docshare.wps.com/document/phone-segmentation-and-recognition-through-phonological-activation-mapping/82149/",4,{"url":50,"name":13,"@type":53,"author":54,"headline":13,"publisher":56,"fileFormat":59,"inLanguage":22,"description":14,"dateModified":60,"datePublished":61,"encodingFormat":59,"isAccessibleForFree":62,"interactionStatistic":63},"DigitalDocument",{"name":9,"@type":55},"Person",{"url":39,"name":57,"@type":58},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":64,"interactionType":65,"userInteractionCount":20},"InteractionCounter",{"@type":66},"ViewAction",{"@type":68,"mainEntity":69},"FAQPage",[70,76,80],{"name":71,"@type":72,"acceptedAnswer":73},"Why is phone segmentation and recognition often modeled separately in modern systems?","Question",{"text":74,"@type":75},"Many approaches frame recognition and segmentation as different task types, such as sequence-to-sequence for recognition and frame-wise boundary classification for segmentation.","Answer",{"name":77,"@type":72,"acceptedAnswer":78},"What is SPAM and what does it produce from S3M representations?",{"text":79,"@type":75},"SPAM maps each S3M representation frame to activations over phonological feature vectors (e.g., voicing, nasality), creating a time-aligned phonological activation map across frames.",{"name":81,"@type":72,"acceptedAnswer":82},"How does the method achieve segmentation and recognition with limited supervision?",{"text":83,"@type":75},"It introduces lightweight recognition and segmentation heads on top of SPAM, requiring less than a minute of phonetic transcriptions and generalizing to unseen phones during training.","https://schema.org",{"og:url":50,"og:type":86,"og:title":13,"og:site_name":57,"og:description":14},"article",{"robots":88,"canonical":50},"index,follow",{"doc_id":7,"site_id":23},{"code":4,"msg":5,"data":91},[92,96,100,104,109,114,119,122,126,129,133],{"id":20,"doc_module":4,"doc_module_name":44,"category_name":93,"show_sort_weight":94,"slug":95},"Story & Novel",90,"story-novel",{"id":45,"doc_module":4,"doc_module_name":44,"category_name":97,"show_sort_weight":98,"slug":99},"Literature",80,"literature",{"id":51,"doc_module":4,"doc_module_name":44,"category_name":101,"show_sort_weight":102,"slug":103},"Exam",70,"exam",{"id":105,"doc_module":4,"doc_module_name":44,"category_name":106,"show_sort_weight":107,"slug":108},5,"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":44,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":44,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":44,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":44,"category_name":124,"show_sort_weight":27,"slug":125},9,"Religion & Spirituality","religion-spirituality",{"id":27,"doc_module":4,"doc_module_name":44,"category_name":127,"show_sort_weight":27,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":44,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":44,"category_name":135,"show_sort_weight":105,"slug":136},19,"General","general"]