[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84990-en":3,"doc-seo-84990-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},84990,7971461740909,"Levi","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","Transformer-based segmentation of prosodic boundaries in Brazilian Portuguese","Automatic prosodic segmentation identifies boundary locations between speech units using acoustic and linguistic evidence. While deep learning methods perform strongly for English, Brazilian Portuguese (BP) segmentation often still depends on rule-based or traditional machine-learning pipelines. The paper introduces SAMPA, a Whisper-based segmenter that transcribes BP and inserts explicit markers for terminal prosodic boundaries. It fine-tunes Whisper large-v3 on manually segmented NURC-SP recordings, tests multiple training and filtering setups, and evaluates out-of-distribution on MuPe-Diversidades. Best results reach F1=0.731 on the held-out test split and F1=0.796 on MuPe-Diversidades, with n-gram and acoustic-visual analyses linking performance to morphosyntactic, semantic, and prosodic cues.","Transformer-based segmentation of prosodic boundaries in Brazilian Portuguese  \n1st Rodrigo de Freitas Lima ICMC University of Sa˜o Paulo So Paulo, Brazil  \n[https://orcid.org/0009-0009-4344-1109](https://orcid.org/0009-0009-4344-1109)  \n2nd Julio Cesar Galdino ICMC University of Sa˜o Paulo So Paulo, Brazil  \n[https://orcid.org/0000-0001-6378-4648](https://orcid.org/0000-0001-6378-4648)  \n3rd Marcos Vinicius Treviso Instituto Superior Te´cnico University of Lisbon Lisbon, Portugal  \n[https://orcid.org/0000-0002-3286-0609](https://orcid.org/0000-0002-3286-0609)  \narXiv :2607 .07408v2 [ cs .CL] 13 Jul 2026  \nAbstract—Automatic prosodic segmentation identifies boundaries between speech units from acoustic and linguistic evidence. Although recent deep learning approaches have produced strong results for English, automatic segmentation for Brazilian Portuguese (BP) still relies mostly on rule-based or traditional machine-learning methods. This paper presents SAMPA, a Whisper-based segmenter that transcribes BP speech while inserting explicit markers for terminal prosodic boundaries. We fine-tune Whisper large-v3 on manually segmented recordings from the NURC-SP dataset and evaluate different training and test-time filtering configurations, including out-of-distribution testing on the MuPe-Diversidades dataset. SAMPA achieves competitive boundary-detection performance across settings, with the best models reaching F1 = 0 .731 on the held-out test split and F1 = 0 .796 on MuPe-Diversidades. Finally, through n-gram and acoustic-visual analyses, we show that our model follows morphosyntactic, semantic, and prosodic cues for detecting prosodic boundaries.  \nIndex Terms—automatic prosodic segmentation, deep learning, Brazilian Portuguese  \nI. INTRODUCTION  \nAutomatic speech segmentation is the task of identifying where boundaries between speech units occur [1] . When these units are delimited by prosodic cues, they are commonly described as intonation units and play several roles in spoken communication, including structuring discourse and organizing information flow [2] .  \nEarlier automatic approaches to prosodic segmentation relied on more traditional machine-learning pipelines. For instance,[3] combined acoustic information with syntactic postprocessing and reported effective boundary detection for English and Russian. More recently, prosodic segmentation has also benefited from large pretrained speech models. [4] introduced PSST!, a Transformer-based approach that fine-tunes Whisper [5] to produce a transcription and mark intonationunit boundaries within the same output sequence.  \nFor Brazilian Portuguese (BP), [6] adapted heuristic methods originally proposed for English [7], relying mainly on  \nThis study was financed, in part, by the So Paulo Research Foundation (FAPESP), Brasil. Proc˜ess Number 2025/23911-6 . This study was financed in  \npart by the Coordenac¸ao de Aperfeic¸oamento de Pessoal de N´ıvel SuperiorBrasil (CAPES) -Finance Code 001 .  \nspeech rate and silent pauses. Such rules are simple and efficient, but their performance may degrade when parameters are not adapted to the target language or when the audio quality is low. [8] used Linear Discriminant Analysis to identify terminal and non-terminal breaks in BP, while [9] proposed a Random Forest classifier based on acoustic features. These studies show that acoustic cues such as fundamental-frequency movement, tessitura changes, and pauses are useful, but they also suggest that some BP intonation units remain difficult to detect with traditional feature-based methods [10] .  \nThis gap motivates the use of large pretrained speech models for BP. Because each language organizes intonation in specific ways, the segmentation strategy that works well for one language cannot be assumed to transfer directly to another [11] . We therefore propose SAMPA (Segmenter for Automatic Marking of Prosodic boundAries in Brazilian Portuguese), a deep learning-based segmenter adapted fro","cbCaicvVor4ssw1Q","https://ap.wps.com/l/cbCaicvVor4ssw1Q","pdf",1470109,2,1,6,"English","en",105,"# Introduction\n# Data","[{\"question\":\"What problem does the paper address in Brazilian Portuguese prosody?\",\"answer\":\"It targets automatic detection of terminal prosodic boundaries between speech units in Brazilian Portuguese, where existing approaches often rely on rules or traditional feature-based methods.\"},{\"question\":\"How does SAMPA perform prosodic boundary segmentation?\",\"answer\":\"SAMPA fine-tunes a Whisper model to transcribe BP speech while inserting explicit markers for terminal prosodic boundaries in the output sequence.\"},{\"question\":\"What datasets and evaluations are used to measure performance?\",\"answer\":\"Training and evaluation use manually segmented BP data from NURC-SP (CORAA NURC-SP Minimal Corpus) and out-of-distribution testing uses MuPe-Diversidades, reporting F1 scores across configurations.\"}]",1784200078,15,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"transformer-based-segmentation-of-prosodic-boundaries-in-brazilian-portuguese","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/transformer-based-segmentation-of-prosodic-boundaries-in-brazilian-portuguese/84990/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-23","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the paper address in Brazilian Portuguese prosody?","Question",{"text":75,"@type":76},"It targets automatic detection of terminal prosodic boundaries between speech units in Brazilian Portuguese, where existing approaches often rely on rules or traditional feature-based methods.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does SAMPA perform prosodic boundary segmentation?",{"text":80,"@type":76},"SAMPA fine-tunes a Whisper model to transcribe BP speech while inserting explicit markers for terminal prosodic boundaries in the output sequence.",{"name":82,"@type":73,"acceptedAnswer":83},"What datasets and evaluations are used to measure performance?",{"text":84,"@type":76},"Training and evaluation use manually segmented BP data from NURC-SP (CORAA NURC-SP Minimal Corpus) and out-of-distribution testing uses MuPe-Diversidades, reporting F1 scores across configurations.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,114,119,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":111,"show_sort_weight":112,"slug":113},"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]