[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-127739-en":3,"doc-seo-127739-105":31,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":30},127739,962084926284,"Aurora","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",6,"Technology","Voice Synthesis Improvement by Machine Learning of Natural Prosody - read online free","Since modern computing emerged, research has pursued seamless human–computer interaction, particularly in computer-generated speech. Many synthetic voices remain identifiable as nonhuman because prosody—especially intonation and rhythm—is often missing or inadequate. This study enhances text-to-speech by introducing machine-learning-based prosodic and paralinguistic elements using an LSTM neural network. A prototype modular platform analyzes and generalizes algorithms, producing more realistic speech and supporting applications in reading, telephony codecs, and improved human–computer interaction.","Edith Cowan University  \nResearch Online  \nResearch outputs 2022 to 2026  \n3-1-2024  \nVoice synthesis improvement by machine learning of natural prosody  \nJoseph Kane  \nEdith Cowan University  \nMichael N. Johnstone Edith Cowan University  \nPatryk Szewczyk Edith Cowan University  \nFollow this and additional works at: [https://ro.ecu.edu.au/ecuworks2022-2026](https://ro.ecu.edu.au/ecuworks2022-2026)  \n Part of the Artificial Intelligence and Robotics Commons  \n10.3390/s24051624  \nKane, J., Johnstone, M. N., & Szewczyk, P. (2024) . Voice synthesis improvement by machine learning of natural prosody. Sensors, 24(5), article 1624. [https://doi.org/10.3390/s24051624](https://doi.org/10.3390/s24051624)  \n[This Journal Article is posted at Research Online.](This Journal Article is posted at Research Online.)[ ](This Journal Article is posted at Research Online.)[https://ro.ecu.edu.au/ecuworks2022-2026/3880](https://ro.ecu.edu.au/ecuworks2022-2026/3880)  \n sensors   \nArticle  \nVoice Synthesis Improvement by Machine Learning of Natural Prosody  \nJoseph Kane 1,2, *,†, Michael N. Johnstone 1,2,† and Patryk Szewczyk 1,2,†  \nCitation: Kane, J.; Johnstone, M.N.; Szewczyk, P. Voice Synthesis Improvement by Machine Learning of Natural Prosody. Sensors 2024, 24, 1624. [https://doi.org/10.3390/](https://doi.org/10.3390/)[ ](https://doi.org/10.3390/)s24051624  \nAcademic Editors: Christoph M. Friedrich and Kit Yan Chan  \nReceived: 3 January 2024  \nRevised: 25 February 2024  \nAccepted: 28 February 2024  \nPublished: 1 March 2024  \nCopyright: © 2024 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license ([https://](https://)[ ](https://)[creativecommons.org/licenses/by/](creativecommons.org/licenses/by/)[ ](creativecommons.org/licenses/by/)[4.0/](4.0/)) .  \n1 Cyber Security Cooperative Research Centre, Edith Cowan University, 270 Joondalup Drive, Joondalup, WA 6027, Australia; [m.johnstone@ecu.edu.au](m.johnstone@ecu.edu.au) (M.N.J.); [p.szewczyk@ecu.edu.au](p.szewczyk@ecu.edu.au) (P.S.)  \n2 Security Research Institute, Edith Cowan University, Joondalup, WA 6027, Australia  \n* [Correspondence: j.kane@ecu.edu.au](Correspondence: j.kane@ecu.edu.au)[ ](Correspondence: j.kane@ecu.edu.au)† These authors contributed equally to this work.  \nAbstract: Since the advent of modern computing, researchers have striven to make the human– computer interface (HCI) as seamless as possible. Progress has been made on various fronts, e.g., the desktop metaphor (interface design) and natural language processing (input) . One area receiving attention recently is voice activation and its corollary, computer-generated speech. Despite decades of research and development, most computer-generated voices remain easily identifiable as nonhuman. Prosody in speech has two primary components—intonation and rhythm—both often lacking in computer-generated voices. This research aims to enhance computer-generated text-to-speech algorithms by incorporating melodic and prosodic elements of human speech. This study explores a novel approach to add prosody by using machine learning, specifically an LSTM neural network, to add paralinguistic elements to a recorded or generated voice. The aim is to increase the realism of computer-generated text-to-speech algorithms, to enhance electronic reading applications, and improved artificial voices for those in need of artificial assistance to speak. A computer that is able to also convey meaning with a spoken audible announcement will also improve human-to-computer interactions. Applications for the use of such an algorithm may include improving high-definition audio codecs for telephony, renewing old recordings, and lowering barriers to the utilization of computing. This research deployed a prototype modular platform for digital speech improvement by analyzing and generalizing algorithms into a modular system through lab","cbCaijWtoqJZ3gkF","https://ap.wps.com/l/cbCaijWtoqJZ3gkF","pdf",1939168,2,1,23,"English","en",105,"# Abstract\n## Introduction\n## Methods and System Design\n## Experimental Evaluation\n## Future Work","[{\"question\":\"Why are computer-generated voices often still perceived as nonhuman?\",\"answer\":\"They typically lack key prosodic components such as intonation and rhythm, which are essential for natural-sounding speech.\"},{\"question\":\"What machine-learning approach does the research use to improve prosody?\",\"answer\":\"The study uses an LSTM neural network to add melodic and prosodic, paralinguistic elements to recorded or generated voice.\"},{\"question\":\"What is the purpose of the prototype system used in the research?\",\"answer\":\"It provides a modular platform that analyzes and generalizes algorithms to optimize combinations and performance, including behavior in edge cases.\"}]","Voice Synthesis Improvement by Machine Learning of Natural Prosody - read online free | PDF",1785941341,58,{"code":4,"msg":32,"data":33},"ok",{"site_id":25,"language":24,"slug":34,"title":13,"keywords":35,"description":14,"schema_data":36,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":29},"voice-synthesis-improvement-by-machine-learning-of-natural-prosody-read-online-free","",{"@graph":37,"@context":86},[38,54,69],{"@type":39,"itemListElement":40},"BreadcrumbList",[41,45,48,51],{"item":42,"name":43,"@type":44,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":46,"name":47,"@type":44,"position":20},"https://docshare.wps.com/document/","Document",{"item":49,"name":12,"@type":44,"position":50},"https://docshare.wps.com/document/technology/",3,{"item":52,"name":13,"@type":44,"position":53},"https://docshare.wps.com/document/voice-synthesis-improvement-by-machine-learning-of-natural-prosody-read-online-free/127739/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":42,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-27","2026-08-05",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"Why are computer-generated voices often still perceived as nonhuman?","Question",{"text":76,"@type":77},"They typically lack key prosodic components such as intonation and rhythm, which are essential for natural-sounding speech.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"What machine-learning approach does the research use to improve prosody?",{"text":81,"@type":77},"The study uses an LSTM neural network to add melodic and prosodic, paralinguistic elements to recorded or generated voice.",{"name":83,"@type":74,"acceptedAnswer":84},"What is the purpose of the prototype system used in the research?",{"text":85,"@type":77},"It provides a modular platform that analyzes and generalizes algorithms to optimize combinations and performance, including behavior in edge cases.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,111,114,119,124,129,132,136],{"id":21,"doc_module":4,"doc_module_name":47,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":47,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":47,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":47,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":11,"doc_module":4,"doc_module_name":47,"category_name":12,"show_sort_weight":112,"slug":113},50,"technology",{"id":115,"doc_module":4,"doc_module_name":47,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":120,"doc_module":4,"doc_module_name":47,"category_name":121,"show_sort_weight":122,"slug":123},8,"Research & Report",30,"research-report",{"id":125,"doc_module":4,"doc_module_name":47,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":47,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":47,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":47,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]