[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-123909-en":3,"doc-seo-123909-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},123909,4398048950312,"Violet","https://ap-avatar.wpscdn.com/avatar/400002538284de19e3c?_k=1778320343897328908",8,"Research & Report","The NeurIPS 2023 Machine Learning for Audio Workshop - Affective Audio Benchmarks and Novel Data","The NeurIPS 2023 Machine Learning for Audio Workshop convenes machine learning experts across audio-relevant domains, addressing the community’s limited accessibility to large, high-quality datasets for time-dependent modalities. It highlights tasks spanning speech emotion recognition, audio event detection, and broader audio generation and modeling needs. To support researchers facing data constraints, the organizers release multiple open-source datasets and provide proprietary datasets during the workshop. The paper describes four large-scale resources—HUME-PROSODY, HUME-VOCALBURST, MODULATE-SONATA, and MODULATE-STREAM—summarizes current baselines, and encourages evaluation beyond the initial benchmark tasks.","arXiv :2403 . 14048v1 [ cs . SD] 21 Mar 2024  \nThe NeurIPS 2023 Machine Learning for Audio Workshop: Affective Audio Benchmarksand Novel Data  \nAlice Baird* Hume AI New York, USA [alice@hume.ai](alice@hume.ai)  \nRachel Manzelli* Modulate AI Massachusetts, USA [rachel@modulate.ai](rachel@modulate.ai)  \nPanagiotis Tzirakis Hume AI New York, USA  \nChris Gagne Hume AI New York, USA  \nHaoqi Li Hume AI New York, USA  \nBrian Kulis Boston University Boston, USA  \nSadie Allen Boston University Boston, USA  \nShrikanth S. Narayanan USC-SAIL California, USA  \nSander Dieleman DeepMind London, UK  \nAlan Cowen Hume AI New York, USA  \nAbstract  \nThe NeurIPS 2023 Machine Learning for Audio Workshop brings together machine learning (ML) experts from various audio domains. There are several valuable audio-driven ML tasks, from speech emotion recognition to audio event detection, but the community is sparse compared to other ML areas, e.g., computer vision or natural language processing. A major limitation with audio is the available data; with audio being a time-dependent modality, high-quality data collection is time-consuming and costly, making it challenging for academic groups to apply their often state-of-the-art strategies to a larger, more generalizable dataset.  \nIn this short white paper, to encourage researchers with limited access to largedatasets, the organizers 􀀂rst outline several open-source datasets that are available to the community, and for the duration of the workshop are making several propriety datasets available. Namely, three vocal datasets, HUME-PROSODY, HUMEVOCALBURST, an acted emotional speech dataset MODULATE-SONATA, and an in-game streamer dataset MODULATE-STREAM. We outline the current baselineson these datasets but encourage researchers from across audio to utilize them outside of the initial baseline tasks.  \n1 Introduction  \nWorking with audio data in machine learning presents unique challenges compared to 􀀂elds like computer vision. Despite the importance of various key problems in the audio domain, such as textto-speech, voice recognition, source separation, and synthesis, it has received considerably lower attention. However, there has been a recent renaissance in audio research, particularly in the 􀀂eld of synthesis, with the release of several in􀀃uential papers in the last year [1–6] . The relative scarcity of prior research and this recent boom serves as the primary motivations behind organizing the 2023 NeurIPS Machine Learning for Audio (MLA) Workshop.  \nThe workshop covers a vast scope of audio-related tasks, including but not limited to speech modeling, speech generation, music generation, denoising of speech and music, data augmentation, acoustic event classi􀀂cation, transcription, source separation, and even multimodal modelling involving audio.  \n*These authors contributed equally to this work.  \nThere are prestigious competitions like the Detection and Classi􀀂cation of Acoustic Scenes and Events (DCASE) focusing on audio-driven machine learning [7], and numerous large-scale audio datasets with high-level labeling available in the literature, including Audioset [8], Urban80k [9], and Librispeech [10], along with unique concepts such as bird identi􀀂cation [11], and music genre detection [12] . Despite the availability of such datasets, there still exists a scarcity of openly accessible large-scale datasets particularly tailored for more specialized domains, such as human-computer interaction and human behavior analysis. The lack of specialized datasets presents a considerable hurdle for developing systems that understand more nuanced human characteristics.  \nTo address this issue and to support authors submitting their work to MLA, the organizers have proactively provided four datasets to researchers participating in the workshop, namely, the HUME-PROSODY, HUME-VOCALBURST, MODULATE-STREAM and MODULATE-SONATA. These datasets offer a broad spectrum of human states, including both labeled and unlabeled d","cbCaivIEm48WWNOm","https://ap.wps.com/l/cbCaivIEm48WWNOm","pdf",179367,1,14,"English","en",105,"# Introduction\n## Workshop scope and motivation\n## Provided datasets and benchmark usage\n# Workshop Audio Datasets\n## Dataset design for speech emotion and human responses\n## HUME-PROSODY\n## HUME-VOCALBURST\n## MODULATE-SONATA\n## MODULATE-STREAM","[{\"question\":\"Why does the workshop emphasize audio datasets and benchmarks?\",\"answer\":\"Audio tasks are constrained by time-dependent, expensive data collection, leaving the community with fewer openly accessible large-scale datasets. The workshop aims to bridge this gap with new resources and established baseline tasks.\"},{\"question\":\"Which four datasets are provided for workshop research?\",\"answer\":\"The paper presents HUME-PROSODY, HUME-VOCALBURST, MODULATE-SONATA, and MODULATE-STREAM. Together they cover human states using different interaction settings and include both labeled and unlabeled data where applicable.\"},{\"question\":\"What do the datasets enable researchers to study?\",\"answer\":\"They support evaluation and improvement of unsupervised models, particularly for speech emotion recognition and modeling human responses to dynamic events. The datasets also provide labeled benchmarks focused on prosody and vocal bursts when applicable.\"}]","The NeurIPS 2023 Machine Learning for Audio Workshop - Affective Audio Benchmarks and Novel Data | PDF",1785819191,35,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"the-neurips-2023-machine-learning-for-audio-workshop-affective-audio-benchmarks-and-novel-data","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/the-neurips-2023-machine-learning-for-audio-workshop-affective-audio-benchmarks-and-novel-data/123909/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why does the workshop emphasize audio datasets and benchmarks?","Question",{"text":75,"@type":76},"Audio tasks are constrained by time-dependent, expensive data collection, leaving the community with fewer openly accessible large-scale datasets. The workshop aims to bridge this gap with new resources and established baseline tasks.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Which four datasets are provided for workshop research?",{"text":80,"@type":76},"The paper presents HUME-PROSODY, HUME-VOCALBURST, MODULATE-SONATA, and MODULATE-STREAM. Together they cover human states using different interaction settings and include both labeled and unlabeled data where applicable.",{"name":82,"@type":73,"acceptedAnswer":83},"What do the datasets enable researchers to study?",{"text":84,"@type":76},"They support evaluation and improvement of unsupervised models, particularly for speech emotion recognition and modeling human responses to dynamic events. The datasets also provide labeled benchmarks focused on prosody and vocal bursts when applicable.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]