[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-146517-en":3,"doc-seo-146517-105":30,"detail-sidebar-cat-0-en-105":90},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},146517,962084925782,"Chloe Bennett","https://ap-avatar.wpscdn.com/davatar_9964176cb1d06d4a9deccf72a44ae3dc",6,"Technology","Unleash Your Inner Monster - A Demonstration of High-Fidelity Human to Non-Human Voice Conversion","Non-human monsters in massive multiplayer online roleplaying games are vital to immersive worlds and strong player engagement, yet high-quality monster audio requires substantial time and budget. AI voice conversion can help, but existing systems often depend on standard human voice training and human-centric audio features, restricting realistic non-human generation. This study proposes a human-to-non-human (H2NH) voice conversion model that captures human speech and converts it in realtime into selected monster vocalizations. The approach records participant voices and produces high-quality non-human sounds suitable for game monster audio pipelines.","Interspeech 2025  \n17-21 August 2025, Rotterdam, The Netherlands  \nUnleash Your Inner Monster:  \nA Demonstration of High-Fidelity Human to Non-Human Voice Conversion  \nNamhyun Cho 1 ,2 ,∗, Sunmin Kim 1 ,∗, Minsu Kang 1, Seolhee Lee 1, Choonghyeon Lee 1, Yangsun Lee 1 ,3  \n1NC AI Co., Ltd, Republic of Korea  \n2 Sogang University, Republic of Korea  \n3 Chung-ang University, Republic of Korea  \n[cnh2769@ncsoft.com](cnh2769@ncsoft.com) , [essems79@ncsoft.com](essems79@ncsoft.com) , [mskang@ncsoft.com](mskang@ncsoft.com) , [seolhee@ncsoft.com](seolhee@ncsoft.com) ,  \n[choonghyeon@ncsoft.com](choonghyeon@ncsoft.com) , [yyslee0301@ncsoft.com](yyslee0301@ncsoft.com)  \nAbstract  \nNon-human monsters in massive multiplayer online roleplaying games (MMORPGs) are essential for immersive game environments and significantly enhance player engagement. However, producing high-quality monster sounds demands considerable time and financial resources. AI-driven voice conversion offers a potential solution, but existing models rely on standard human voice training and human-centric audio features, limiting their ability to generate realistic non-human voices. This study presents a human-to-non-human (H2NH) voice conversion model designed to address these challenges. The model effectively generated high-quality non-human sounds by recording the voices of participants and converting them in realtime into selected monster vocalizations.  \nIndex Terms: non-human voice conversion, monster sound  \n1. Introduction  \nThe global gaming market expanded rapidly during the COVID- 19 pandemic as demand for stay-at-home, non-face-to-face entertainment increased [1] . Recent massive multiplayer online role-playing games (MMORPGs), such as Lineage W, now feature high-quality 3D graphics and 48 kHz high-fidelity sound to improve realism. Medieval fantasy–themed games, inspired by Dungeons & Dragons [2], often incorporate fictional monsters such as dragons and orcs to enrich the virtual world. Achieving such immersive environments requires extensive creativity not only from graphic designers shaping a monster’s appearance but also from sound designers crafting its voice—an effort that incurs significant costs. This study proposes a voice conversion model that transforms human voices into non-human timbres, reducing production costs. Section 2 examines the characteristics of monster sounds. Section 3 reviews recent voice conversion technologies and their limitations. Section 4 introduces the proposed model, while Section 5 details the implementation and demonstration of the model. Section 6 presents the conclusions.  \n2. The Monster Sound  \nMonster sounds in games occupy a unique position between voice and sound effects. While they convey linguistic information, making them voice-like, they also require extensive filtering to achieve textures that define the identity of the monster, classifying them as a specialized type of sound effect. Additionally, they incorporate elements rarely addressed by conventional speech technology, such as unvoiced sound (e.g., breathing) and extreme vocal expressions like screaming. Therefore, producing high-quality monster voices demands exceptional performances from voice actors and extensive postproduction work,  \n* Co-authors.  \n(a) Source audio of ”Born a slave, died a king.” performed in Korean.  \n(b) Result audio of voice conversion into an Orc voice.  \nFigure 1: Spectrogram comparison of a human voice and an Orc-converted monster sound.  \nresulting in higher time and financial costs than other in-game audios.  \nThe production process generally involves three stages: (1) defining the timbre of the target monster, (2) hiring a professional voice actor capable of approximating that timbre, and (3) applying effects and modulation to refine the recorded voice. The timbre of a monster is largely determined by its origin, size, and sex. Employing skilled voice actors is costly because the vocal range required is broader than standard ","cbCaihUB4auzQXbH","https://ap.wps.com/l/cbCaihUB4auzQXbH","pdf",1143763,1,2,"English","en",105,"# Introduction\n# The Monster Sound\n# Related Works\n# Proposed Method\n# Implementation and Demonstration\n# Conclusions","[{\"question\":\"Why is generating monster sounds costly in games?\",\"answer\":\"Monster sounds sit between speech and sound effects and require extensive filtering, effects, and postproduction. Achieving convincing monster timbre also demands skilled voice actors and specialized tools.\"},{\"question\":\"What limitation of existing voice conversion models affects non-human voice generation?\",\"answer\":\"Most models rely on speech features and training data that represent human vocal characteristics. They struggle to reproduce timbre for reference monster voices with broader frequency ranges and intricate audio effects.\"},{\"question\":\"How does the proposed H2NH voice conversion model work?\",\"answer\":\"The model uses Mel spectrograms and a reference encoder to extract global style. It combines that with self-supervised perturbations to capture linguistic features, then converts human voices in realtime into chosen monster vocalizations.\"}]","Unleash Your Inner Monster - A Demonstration of High-Fidelity Human to Non-Human Voice Conversion | PDF",1787744093,5,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":85,"head_meta":87,"extra_data":89,"updated_unix":28},"unleash-your-inner-monster-a-demonstration-of-high-fidelity-human-to-non-human-voice-conversion","",{"@graph":36,"@context":84},[37,53,67],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":21},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/technology/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/unleash-your-inner-monster-a-demonstration-of-high-fidelity-human-to-non-human-voice-conversion/146517/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":61,"encodingFormat":60,"isAccessibleForFree":62,"interactionStatistic":63},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-08-26",true,{"@type":64,"interactionType":65,"userInteractionCount":4},"InteractionCounter",{"@type":66},"ViewAction",{"@type":68,"mainEntity":69},"FAQPage",[70,76,80],{"name":71,"@type":72,"acceptedAnswer":73},"Why is generating monster sounds costly in games?","Question",{"text":74,"@type":75},"Monster sounds sit between speech and sound effects and require extensive filtering, effects, and postproduction. Achieving convincing monster timbre also demands skilled voice actors and specialized tools.","Answer",{"name":77,"@type":72,"acceptedAnswer":78},"What limitation of existing voice conversion models affects non-human voice generation?",{"text":79,"@type":75},"Most models rely on speech features and training data that represent human vocal characteristics. They struggle to reproduce timbre for reference monster voices with broader frequency ranges and intricate audio effects.",{"name":81,"@type":72,"acceptedAnswer":82},"How does the proposed H2NH voice conversion model work?",{"text":83,"@type":75},"The model uses Mel spectrograms and a reference encoder to extract global style. It combines that with self-supervised perturbations to capture linguistic features, then converts human voices in realtime into chosen monster vocalizations.","https://schema.org",{"og:url":51,"og:type":86,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":88,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":91},[92,96,100,104,108,111,116,121,126,129,133],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":93,"show_sort_weight":94,"slug":95},"Story & Novel",90,"story-novel",{"id":21,"doc_module":4,"doc_module_name":46,"category_name":97,"show_sort_weight":98,"slug":99},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":101,"show_sort_weight":102,"slug":103},"Exam",70,"exam",{"id":29,"doc_module":4,"doc_module_name":46,"category_name":105,"show_sort_weight":106,"slug":107},"Comic",60,"comic",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":109,"slug":110},50,"technology",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},7,"Healthcare",40,"healthcare",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":118,"show_sort_weight":119,"slug":120},8,"Research & Report",30,"research-report",{"id":122,"doc_module":4,"doc_module_name":46,"category_name":123,"show_sort_weight":124,"slug":125},9,"Religion & Spirituality",20,"religion-spirituality",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":127,"show_sort_weight":124,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":46,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":46,"category_name":135,"show_sort_weight":29,"slug":136},19,"General","general"]