[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-133983-en":3,"doc-seo-133983-105":31,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":30},133983,687207412472,"Angel","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","Designed Vocalizations Dataset - Sound-Designed Human and Animal Voices for Non-human Voice Conversion - Dataset and Benchmark Paper","Advances in AI-based voice conversion have expanded media applications such as films, audiobooks, and games, yet most benchmarks remain centered on natural human speech. Designed Vocalizations Dataset addresses the gap by curating diverse raw vocal sources, including speech and animal vocalizations, and applying professional vocal effects processing to create effect-modified variants. A standardized test set provides seen/unseen splits across source timbre groups and preset styles, enabling controlled generalization evaluation. Baseline benchmark results and demo samples support reproducible research.","Designed Vocalizations Dataset: Sound-Designed Human and Animal Voices  \nfor Non-human Voice Conversion  \nSeolhee Lee 1, Minsu Kang 1, Yangsun Lee 1, Woosun Min 1, Choonghyeon Lee 1, Namhyun Cho 1 ,2  \n1 NC AI Co., Ltd, Republic of Korea  \n2 Sogang University, Republic of Korea  \n{seolhee, mskang, yyslee0301, choonghyeon, [cnh2769](cnh2769}@ncsoft.com)[}](cnh2769}@ncsoft.com)[@ncsoft.com](cnh2769}@ncsoft.com) , [woosun1108@naver.com](woosun1108@naver.com)  \narXiv :2607 .20951v1 [ ee ss .AS] 23 Jul 2026  \nAbstract  \nAdvances in AI-based voice conversion have enabled a wide range of media applications, including films, audiobooks, and games. However, most research and public benchmarks still focus on natural human speech, leaving designed vocalizations, such as monster growls and robotic voices, underexplored, partly due to the lack of publicly available resources. To address this gap, we introduce the Designed Vocalizations Dataset, constructed by curating diverse raw vocal sources, including speech and animal vocalizations, and applying professional vocal effects processing to produce corresponding effectmodified variants. We further provide a standardized test set with explicit seen/unseen splits over source timbre groups and preset styles to assess generalization under controlled conditions. Finally, we report baseline benchmark results to support reproducible evaluation and future research. The dataset and demo samples are available online. 1  \nIndex Terms: dataset, designed vocalizations, non-human voice conversion, style/timbre transfer, automated sound design  \n1. Introduction  \nThe growth of interactive and creative media industries (games, film, animation, VR/AR, etc.) has driven increasing demand for diverse vocal sounds that enhance character expression and immersion. In particular, non-natural/non-human vocalizations such as monster roars, robotic voices, and stylized character utterances are essential for shaping a content’s identity and narrative atmosphere. In production, these vocalizations are difficult to capture through recording alone. Sound designers typically rely on complex DSP chains, including distortion, spectral transformation, modulation, and multitrack mixing, which require iterative construction and careful tuning. Consequently, producing high-quality designed vocalizations remains a laborintensive manual workflow, motivating the development of automated or assistive tools.  \nRecent work has proposed human-to-non-human voice conversion (H2NH-VC) methods targeting non-natural, non-human timbres [1, 2, 3, 4] . However, most of these efforts rely on internally curated datasets and evaluation resources, which are rarely released publicly. Data compositions and evaluation protocols also differ across studies, making fair comparison under matched conditions difficult. This lack of open resources contrasts with natural speech research, where dedicated public datasets [5, 6, 7, 8] and benchmarks have enabled reproducible evaluation and accelerated progress across TTS and voice conversion [9, 10, 11, 12, 13, 14, 15, 16] . To support system-  \n1 [https://ncai-official.github.io/speech/](https://ncai-official.github.io/speech/)[ ](https://ncai-official.github.io/speech/)publications/designed-vocalizations-dataset/  \natic comparison and generalization tests in the non-human domain, publicly available datasets and standardized benchmarks are still needed.  \nTo address this gap, we introduce the Designed Vocalizations Dataset, a public dataset of sound-designed vocalizations. The dataset covers linguistic sources and diverse non-linguistic vocalizations, including animal sounds, interjections, and vocal mimicry. In addition to raw recordings, it includes designed vocalizations used in real-world sound design. The designed vocalizations are produced by professional sound designers using effect-chain presets, spanning styles such as monster and creature vocalizations, robotic voices, low-register power voices, and","cbCailX95PAYfmDn","https://ap.wps.com/l/cbCailX95PAYfmDn","pdf",476190,3,1,5,"English","en",105,"# Introduction\n## Motivation and gap in public resources\n## Dataset overview and sound-designed vocalizations\n## Standardized test set and evaluation protocol\n## Benchmark baselines and contributions","[{\"question\":\"Why is a designed vocalizations dataset needed for non-human voice conversion?\",\"answer\":\"Most existing research and public benchmarks focus on natural human speech, while designed vocalizations (e.g., monster growls and robotic voices) are underexplored due to limited publicly available resources.\"},{\"question\":\"How is the Designed Vocalizations Dataset constructed?\",\"answer\":\"The dataset curates diverse raw vocal sources (speech and animal vocalizations) and applies professional vocal effects processing to generate effect-modified variants, including real-world sound design outputs.\"},{\"question\":\"What evaluation setting does the dataset provide for generalization?\",\"answer\":\"It includes a standardized test set with explicit seen/unseen splits across preset styles and source timbre groups, allowing controlled assessment of generalization to unseen styles and unseen sources.\"}]","Designed Vocalizations Dataset - Sound-Designed Human and Animal Voices for Non-human Voice Conversion - Dataset and Benchmark Paper | PDF",1787230963,13,{"code":4,"msg":32,"data":33},"ok",{"site_id":25,"language":24,"slug":34,"title":13,"keywords":35,"description":14,"schema_data":36,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":29},"designed-vocalizations-dataset-sound-designed-human-and-animal-voices-for-non-human-voice-conversion-dataset-and-benchmark-paper","",{"@graph":37,"@context":86},[38,54,69],{"@type":39,"itemListElement":40},"BreadcrumbList",[41,45,49,51],{"item":42,"name":43,"@type":44,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":46,"name":47,"@type":44,"position":48},"https://docshare.wps.com/document/","Document",2,{"item":50,"name":12,"@type":44,"position":20},"https://docshare.wps.com/document/research-report/",{"item":52,"name":13,"@type":44,"position":53},"https://docshare.wps.com/document/designed-vocalizations-dataset-sound-designed-human-and-animal-voices-for-non-human-voice-conversion-dataset-and-benchmark-paper/133983/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":42,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-09-01","2026-08-20",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"Why is a designed vocalizations dataset needed for non-human voice conversion?","Question",{"text":76,"@type":77},"Most existing research and public benchmarks focus on natural human speech, while designed vocalizations (e.g., monster growls and robotic voices) are underexplored due to limited publicly available resources.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How is the Designed Vocalizations Dataset constructed?",{"text":81,"@type":77},"The dataset curates diverse raw vocal sources (speech and animal vocalizations) and applies professional vocal effects processing to generate effect-modified variants, including real-world sound design outputs.",{"name":83,"@type":74,"acceptedAnswer":84},"What evaluation setting does the dataset provide for generalization?",{"text":85,"@type":77},"It includes a standardized test set with explicit seen/unseen splits across preset styles and source timbre groups, allowing controlled assessment of generalization to unseen styles and unseen sources.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":47,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":48,"doc_module":4,"doc_module_name":47,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":47,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":22,"doc_module":4,"doc_module_name":47,"category_name":107,"show_sort_weight":108,"slug":109},"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":47,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":47,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":47,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":47,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":47,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":47,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":47,"category_name":137,"show_sort_weight":22,"slug":138},19,"General","general"]