[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83867-en":3,"doc-seo-83867-105":30,"detail-sidebar-cat-0-en-105":90},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83867,8796095462418,"Noah","https://ap-avatar.wpscdn.com/avatar/80000253c1241d02b47?x-image-process=image/resize,m_fixed,w_180,h_180&k=1778826106357471780",8,"Research & Report","DuplexChat: Constructing Speaker-Separated Full-Duplex Dialogue Speech at Scale for Spoken Dialogue Language Modeling","Full-duplex spoken dialogue models require conversations recorded as separate speaker tracks to preserve turn-taking, backchannels, and overlaps, yet most large public speech corpora are monaural and unsuitable for SDLM training. DuplexChat introduces an open-source corpus and DuplexChat-Pipe pipeline to build speaker-separated full-duplex dialogue from public podcast feeds. The pipeline filters feeds by language, retrieves and cleans audio, uses diarization-guided segmentation, and applies speech separation and restoration. Results show 282,634 hours of English and 132,723 hours of Japanese, with turn-taking dynamics matching human dialogue.","DuplexChat: Constructing Speaker-Separated Full-Duplex Dialogue Speech at Scale for Spoken Dialogue Language Modeling  \nWataru Nakata 1 ,2 , Yuki Saito 1 ,2 and Hiroshi Saruwatari 1  \n1The University of Tokyo, Japan, 2National Institute of Advanced Industrial Science and Technology, Japan.  \n{nakata-wataru855, [sythonuk](sythonuk}@g.ecc.u-tokyo.ac.jp)[}](sythonuk}@g.ecc.u-tokyo.ac.jp)[@g.ecc.u-tokyo.ac.jp](sythonuk}@g.ecc.u-tokyo.ac.jp)  \narXiv :2607 .0494 1v 1 [ cs .CL] 6 Jul 2026  \nAbstract—Full-duplex spoken dialogue models are trained on conversational speech in which each speaker is represented as a separate stream, but existing large-scale public speech corpora are mostly monaural, making them unsuited for SDLM training. We present DuplexChat, an open-source corpus for fullduplex spoken dialogue models, and DuplexChat-Pipe, a pipeline for constructing speaker-separated full-duplex dialogue speech from public podcast feeds. DuplexChat-Pipe filters languagespecific podcast feeds, retrieves and cleans episode audio, extracts diarization-guided two-speaker dialogue clips, and applies speech separation and restoration to produce one channel per speaker. Running this pipeline yields a speaker-separated spoken dialogue corpus covering 282,634 hours of English and 132,723 hours of Japanese. Analysis results on DuplexChat show that it contains turn-taking dynamics present in human dialogues.  \nI. INTRODUCTION  \nSpeech is a primary medium of human communication and a convenient interface for conversational AI. However, natural spoken interaction requires more than converting text responses into speech. Human conversation is tightly coordinated in real time. Speakers take turns with short gaps, provide backchannels, and frequently overlap [1] . These conversational dynamics carry essential information that enables rich, responsive interaction, yet they are difficult to capture with conventional cascaded spoken dialogue systems [2], which serialize automatic speech recognition, text generation, and speech synthesis [3] . Spoken Dialogue Language Models (SDLMs) such as dGSLM [4] and Moshi [5] address this limitation by modeling the dialogue in an end-to-end manner, making them a promising direction for more natural voice interaction.  \nAs with text language models [6], SDLMs are expected to benefit from larger training corpora, yet the data they require is difficult to obtain at scale. Training requires fullduplex dialogue speech [4], [5]: conversations in which the two participants are recorded as separate audio tracks, so that turn-taking, backchannels, and overlapping speech remain observable. Existing full-duplex resources are dominated by telephone-conversation corpora such as CallHome [7],[8], [9],[10],[11] and Fisher [12], whose collection requires recruiting paired participants and therefore does not scale easily. In contrast, large public speech corpora such as GigaSpeech [13], YODAS [14], and the Spotify Podcast Dataset [15] provide web-scale audio but consist of monaural recordings. Once two  \nspeakers have been mixed into a single channel, the speakerspecific timing needed for full-duplex modeling is no longer directly available. What is missing is an open and scalable method that combines web-scale audio with speaker-separated channels.  \nScaling such a corpus beyond recorded corpora calls for harvesting audio from the Internet rather than collecting telephone conversations. Podcasts are a natural source to crawl: distributed through public RSS feeds, they supply an effectively unlimited and continually growing stream of spontaneous dialogue. However, podcast audio is mixed-channel audio and interleaved with substantial non-dialogue material such as monologues, advertisements, and music.  \nIn this paper, we propose DuplexChat-Pipe, an open-source pipeline that constructs full-duplex dialogue speech by crawling public podcast feeds, and the resulting corpus named DuplexChat. DuplexChat is a spoken dialogue collection comp","cbCaib6ozUYv4jEG","https://ap.wps.com/l/cbCaib6ozUYv4jEG","pdf",246406,3,1,4,"English","en",105,"# Introduction\n# DuplexChat-Pipe\n## Feed collection\n## Audio retrieval and cleaning\n## Diarization-based dialogue segmentation\n## Speech separation and restoration","[{\"question\":\"Why are existing large-scale public speech corpora unsuitable for full-duplex spoken dialogue language model (SDLM) training?\",\"answer\":\"Most public corpora are monaural, so mixing two speakers into a single channel removes speaker-specific timing needed to model turn-taking, backchannels, and overlapping speech.\"},{\"question\":\"What are DuplexChat and DuplexChat-Pipe?\",\"answer\":\"DuplexChat is an open-source speaker-separated full-duplex spoken dialogue corpus, and DuplexChat-Pipe is the open-source pipeline that constructs it from public podcast feeds.\"},{\"question\":\"How does DuplexChat-Pipe generate one channel per speaker from podcast audio?\",\"answer\":\"It filters language-specific feeds, retrieves and cleans episode audio, extracts two-speaker dialogue clips using diarization guidance, then applies speech separation and restoration to produce speaker-separated channels.\"}]",1784191091,10,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":85,"head_meta":87,"extra_data":89,"updated_unix":28},"duplexchat-constructing-speaker-separated-full-duplex-dialogue-speech-at-scale-for-spoken-dialogue-language-modeling","",{"@graph":36,"@context":84},[37,52,67],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":22},"https://docshare.wps.com/document/duplexchat-constructing-speaker-separated-full-duplex-dialogue-speech-at-scale-for-spoken-dialogue-language-modeling/83867/",{"url":51,"name":13,"@type":53,"author":54,"headline":13,"publisher":56,"fileFormat":59,"inLanguage":24,"description":14,"dateModified":60,"datePublished":61,"encodingFormat":59,"isAccessibleForFree":62,"interactionStatistic":63},"DigitalDocument",{"name":9,"@type":55},"Person",{"url":41,"name":57,"@type":58},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":64,"interactionType":65,"userInteractionCount":20},"InteractionCounter",{"@type":66},"ViewAction",{"@type":68,"mainEntity":69},"FAQPage",[70,76,80],{"name":71,"@type":72,"acceptedAnswer":73},"Why are existing large-scale public speech corpora unsuitable for full-duplex spoken dialogue language model (SDLM) training?","Question",{"text":74,"@type":75},"Most public corpora are monaural, so mixing two speakers into a single channel removes speaker-specific timing needed to model turn-taking, backchannels, and overlapping speech.","Answer",{"name":77,"@type":72,"acceptedAnswer":78},"What are DuplexChat and DuplexChat-Pipe?",{"text":79,"@type":75},"DuplexChat is an open-source speaker-separated full-duplex spoken dialogue corpus, and DuplexChat-Pipe is the open-source pipeline that constructs it from public podcast feeds.",{"name":81,"@type":72,"acceptedAnswer":82},"How does DuplexChat-Pipe generate one channel per speaker from podcast audio?",{"text":83,"@type":75},"It filters language-specific feeds, retrieves and cleans episode audio, extracts two-speaker dialogue clips using diarization guidance, then applies speech separation and restoration to produce speaker-separated channels.","https://schema.org",{"og:url":51,"og:type":86,"og:title":13,"og:site_name":57,"og:description":14},"article",{"robots":88,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":91},[92,96,100,104,109,114,119,122,127,130,133],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":93,"show_sort_weight":94,"slug":95},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":97,"show_sort_weight":98,"slug":99},"Literature",80,"literature",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":101,"show_sort_weight":102,"slug":103},"Exam",70,"exam",{"id":105,"doc_module":4,"doc_module_name":46,"category_name":106,"show_sort_weight":107,"slug":108},5,"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":46,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":29,"doc_module":4,"doc_module_name":46,"category_name":131,"show_sort_weight":29,"slug":132},"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":46,"category_name":135,"show_sort_weight":105,"slug":136},19,"General","general"]