[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85080-en":3,"doc-seo-85080-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85080,1099514067438,"River Wang","https://ap-avatar.wpscdn.com/avatar/100002539ee87300030?x-image-process=image/resize,m_fixed,w_180,h_180&k=1780474512215547542",8,"Research & Report","On the Role of Conversational Timing in Synthetic Training Data for ASR","Synthetic multi-speaker conversations are widely used to train conversational automatic speech recognition (ASR), yet it is unclear which timing properties make simulated data genuinely useful. The paper models pause and overlap timing as a controllable training variable using an exponential-tilting family estimated from multiple corpora. Timing configurations are sampled via Latin hypercube sampling and refined with multi-objective Bayesian optimization, then used to train ASR and assess cpWER/cpCER on a Hungarian dialogue corpus. Results link ASR behavior to induced timing statistics and reveal an overlap–gap trade-off for simulated training data.","On the Role of Conversational Timing in Synthetic Training Data for ASR  \nMt Gedeon∗ ,†, Pter Mihajlik∗  \n∗Dept. of Telecommunications and Artificial Intelligence, Budapest University of Technology and Economics, Hungary  \n†Speechtex Ltd.  \n[gedeonm@edu.bme.hu](gedeonm@edu.bme.hu) , [mihajlik@tmit.bme.hu](mihajlik@tmit.bme.hu)  \narXiv :2607 .08371v1 [ ee ss .AS] 9 Jul 2026  \nAbstract—Synthetic multi-speaker conversations are widely used to train conversational automatic speech recognition (ASR) systems, but it remains unclear which timing properties make simulated data most useful. This paper studies conversational timing as a controllable training variable rather than merely as a corpus statistic to be reproduced. We parameterize pause and overlap timing distributions with an exponential-tilting family estimated from multiple conversational corpora, and then explore the resulting four-dimensional parameter space with Latin hypercube sampling and multi-objective Bayesian optimization. Each sampled timing configuration is used to generate simulated training conversations, train an ASR system, and evaluate concatenated-permutation word and character error rates (cpWER and cpCER) on a Hungarian dialogue corpus. The results show that downstream ASR behavior is explained more directly by induced timing statistics than by raw simulator coordinates or corpus proximity. In particular, higher overlap exposure is associated with lower cpWER, whereas longer and more variable gaps are associated with higher cpWER; cpCER follows the same trend, but with weaker statistical support. Bayesian optimization yields modest aggregate improvements, but its main value is analytical: it produces controlled timing interventions that reveal an overlap–gap trade-off in simulated conversational training data. These findings suggest that realistic simulation should be complemented by task-relevant diagnostics of overlap, gap, and timing-variability profiles.  \nIndex Terms—Automatic speech recognition, conversational speech, data simulation, multi-speaker speech, speech data augmentation, overlapped speech, Bayesian optimization.  \nI. INTRODUCTION  \nCONVERSATIONAL speech simulation is widely used in  \nmodern speech technology when natural multi-speaker data are limited in size, diversity, or annotation quality [1]–[3] . A common approach is to construct artificial conversations from single-speaker recordings, making it possible to generate large amounts of training data while controlling interaction properties such as pause duration, overlap frequency, and turn-taking behavior [4]–[6] . These properties are more than descriptive statistics of a conversation: they affect segmentation ambiguity, acoustic interference, speaker attribution, and, ultimately, recognition and diarization performance [7]–[9] . Most existing simulation pipelines are designed around a realism objective. They estimate timing statistics from conversational corpora and generate synthetic interactions that reproduce those observed distributions [5], [6], [10] . This  \nis a natural and useful design principle, because unrealistic timing can produce training data that differ substantially from spontaneous conversations. However, realism alone does not answer a second question that is central to training: which timing properties make simulated data useful for a downstream model? A corpus-faithful distribution may be realistic, but it is not necessarily the most useful distribution for automatic speech recognition (ASR) or end-to-end neural diarization (EEND) .  \nThis distinction motivates the present work. Instead of treating simulation only as a way to imitate a target corpus, we treat conversational timing as a controllable object for systematic analysis. In particular, we ask how overlap rate, mean overlap duration, pause statistics, tail behavior, and the position of a distribution in a predefined parameter space relate to downstream ASR performance. The goal is not simply to identify ","cbCaiqbZn6d2ih4N","https://ap.wps.com/l/cbCaiqbZn6d2ih4N","pdf",480125,5,1,13,"English","en",105,"# Introduction\n## Conversational speech simulation and timing properties\n## Exponential-tilting parameterization\n## Sampling and Bayesian optimization framework\n## Empirical evaluation on BEADialogue","[{\"question\":\"What is the main goal of treating conversational timing as a controllable variable?\",\"answer\":\"The work aims to identify which timing properties of simulated conversations are linked to downstream ASR usefulness, rather than only reproducing corpus statistics for realism.\"},{\"question\":\"How are pause and overlap timing distributions parameterized?\",\"answer\":\"Pause and overlap timing are parameterized using an exponential-tilting family estimated from multiple conversational corpora, forming a low-dimensional parameter space for controlled interventions.\"},{\"question\":\"What timing relationships are reported with ASR error (cpWER/cpCER)?\",\"answer\":\"Higher overlap exposure is associated with lower cpWER, while longer and more variable gaps are associated with higher cpWER; cpCER follows the same direction but with weaker statistical support.\"}]",1784200924,33,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"on-the-role-of-conversational-timing-in-synthetic-training-data-for-asr","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/on-the-role-of-conversational-timing-in-synthetic-training-data-for-asr/85080/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What is the main goal of treating conversational timing as a controllable variable?","Question",{"text":76,"@type":77},"The work aims to identify which timing properties of simulated conversations are linked to downstream ASR usefulness, rather than only reproducing corpus statistics for realism.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How are pause and overlap timing distributions parameterized?",{"text":81,"@type":77},"Pause and overlap timing are parameterized using an exponential-tilting family estimated from multiple conversational corpora, forming a low-dimensional parameter space for controlled interventions.",{"name":83,"@type":74,"acceptedAnswer":84},"What timing relationships are reported with ASR error (cpWER/cpCER)?",{"text":85,"@type":77},"Higher overlap exposure is associated with lower cpWER, while longer and more variable gaps are associated with higher cpWER; cpCER follows the same direction but with weaker statistical support.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":20,"slug":138},19,"General","general"]