[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84711-en":3,"doc-seo-84711-105":28,"detail-sidebar-cat-0-en-105":90},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":11,"language":21,"language_code":22,"site_id":23,"html_lang":22,"table_of_contents":24,"faqs":25,"seo_title":13,"seo_description":14,"update_tm":26,"read_time":27},84711,549758252649,"Ivy","https://ap-avatar.wpscdn.com/avatar/8000253669c5317157?_k=1778319167496531819",8,"Research & Report","A New Method for Identifying Synthetic Speech Fakes based on Information-Geometric Superposed Vowel Evaluation: Part 1. Moraic Syllabary (Japanese)","A method for distinguishing AI-generated fake human speech from natural speech is presented using Japanese as an example. Synthetic speech draws on a limited set of training vowel spectra, yielding less diverse vowel distributions. Natural speech shows richer spectral variety due to flexible human articulation. By analyzing normalized vowel mora spectra as probability densities and measuring distances with the Wasserstein metric, short inter-vowel distances for synthetic speech are preserved under topological mapping with persistent homology, producing distinct clusters.","# A New Method for\n\nIdentifying Synthetic Speech Fakes based onInformation-Geometric Superposed Vowel Evaluation:Part 1.Moraic Syllabary (Japanese)  \nYusei TAMURA  \nGraduate School of Interdisciplinary Information StudiesThe University of Tokyotamura-vusei@g.ecc.u-tokyo.ac.ijp  \nShigekazu ISHIHARA  \nDepartment of Psychology,Faculty of Health and Wellness SciencesHiroshima International University  \ni-shige@hirokoku-u.ac.jp  \nKen ITO  \nInterfaculty Initiative in Information StudiesThe University of Tokyo  \nitosec@iii.u-tokyo.ac.jp  \nJune 30,2026  \n## Abstract\n\nThis paper explains the principles and provides examples of a new method for distinguishing between FAKE humanspeech synthesized by generative AIand natural speech.Since synthetic speech is generated based on information froma limited set of training spectra,the variety of vowels-which are key toidentifying individuals-is limited.In contrast,natural speech exhibits a more diverse distribution of vowel spectra due to the flexibility of the human articulatoryorgan.In this paper,using Japanese-a Syllabary limited to five vowel phonemes,each of which corresponds one-to-one with a specific sound-as an example,we outline a method for distinguishing between synthetic and natural speechreading the same text by analyzing the spectral distributions.If we normalize the spectra of speech sounds and regardthem as probability densityfunctions for the frequency bands received by the hair cells of human cochlea,and evaluatethe distance between spectra using the Wasserstein metric,the Wasserstein distances between the vowels of syntheticspeech are short.By preserving this distance and performing a topological mapping using persistent homology,theespectral probability density functions of synthetic and natural speech can be decomposed into clusters.  \nKeywords:deep fake,speech analysis,Information Geometry,Stochastic Spectroscopy,Wasserstein distance,Graph Laplacian,Persistent homology,Topological mapping,Syllabary  \n## 1 Introduction\n\nDeepfakes have become a social issue and are used invarious crimes,such as investment scams orkidnapping [1][2].According to(Japanese)statistics,90 percent of the public is aware of their existence andthat it is difficult to distinguish them solely bylistening [3].However,only about 20 to 40 percent ofthe public is aware that synthetic speech can be clearlydistinguished using appropriate methods [4],whilethe majority mistakenly believe otherwise.In Japan,these misconceptions are widespread,even amongcabinet ministers and members of the Diet.  \nIn this paper,we explain the principles andprovide practical examples of a new method forclearly distinguishing AI-generated fake humanvoices from natural spoken language,building on ourprevious researchinto musical instrument timbres [5].  \nSynthetic speech learns from a limited set oftraining spectra to imitate the speech of a specificindividualSince it is based on information from afinite set of vowel spectra,its variety is intrinsicallylimited.In contrast,natural speech exhibits muchgreater variety because the flexibility of the humanvocal tract and the redundancy of spoken-languagevowels allow a single vowel to correspond to manydifferent acoustic waveforms.  \nThis paper provides an overview of methods fordetecting deepfakes in Japanese;other languages,such as English,will be discussed in detail insubsequent papers.Japanese is regarded as moraicsyllabary and its vowelletters are limited to five-“A,  \nI,U,E,O”-and each corresponds to a specific moraicsyllable one to one.  \nUsing Japanese example sentences,by extractingthe spectra of each vowel mora in both synthesizedand natural speech as they read aloud the same text,and normalizing the spectra by dividing by theirintegral,it is possible to interpret these spectra asprobability density functions for the frequency bands  \nreceived by human hair cells in the cochlea(probabilistic spectroscopy).  \nWhen the distances between spectra areevaluated using the Wasserstein ","cbCaijZIVJcHk79X","https://ap.wps.com/l/cbCaijZIVJcHk79X","pdf",787618,1,"English","en",105,"# Abstract\n# 1 Introduction\n# 2 Theory","[{\"question\":\"Why is synthetic speech harder to distinguish using only listening?\",\"answer\":\"The paper notes that the public often believes synthetic speech is difficult to detect by listening alone, while only a smaller fraction knows it can be clearly distinguished using appropriate methods.\"},{\"question\":\"What property of Japanese synthetic speech leads to the proposed detection advantage?\",\"answer\":\"Because synthetic speech is learned from a limited set of training spectra, the variety of vowel spectra is intrinsically limited compared with natural speech.\"},{\"question\":\"How does the method separate synthetic from natural speech?\",\"answer\":\"It extracts and normalizes vowel mora spectra from the same read text, treats them as probability density functions, computes Wasserstein distances, and applies persistent homology topological mapping to decompose spectra into distinct clusters for synthetic versus natural speech.\"}]",1784197784,20,{"code":4,"msg":29,"data":30},"ok",{"site_id":23,"language":22,"slug":31,"title":13,"keywords":32,"description":14,"schema_data":33,"social_meta":85,"head_meta":87,"extra_data":89,"updated_unix":26},"a-new-method-for-identifying-synthetic-speech-fakes-based-on-information-geometric-superposed-vowel-evaluation-part-1-moraic-syllabary-japanese","",{"@graph":34,"@context":84},[35,52,67],{"@type":36,"itemListElement":37},"BreadcrumbList",[38,42,46,49],{"item":39,"name":40,"@type":41,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":43,"name":44,"@type":41,"position":45},"https://docshare.wps.com/document/","Document",2,{"item":47,"name":12,"@type":41,"position":48},"https://docshare.wps.com/document/research-report/",3,{"item":50,"name":13,"@type":41,"position":51},"https://docshare.wps.com/document/a-new-method-for-identifying-synthetic-speech-fakes-based-on-information-geometric-superposed-vowel-evaluation-part-1-moraic-syllabary-japanese/84711/",4,{"url":50,"name":13,"@type":53,"author":54,"headline":13,"publisher":56,"fileFormat":59,"inLanguage":22,"description":14,"dateModified":60,"datePublished":61,"encodingFormat":59,"isAccessibleForFree":62,"interactionStatistic":63},"DigitalDocument",{"name":9,"@type":55},"Person",{"url":39,"name":57,"@type":58},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":64,"interactionType":65,"userInteractionCount":20},"InteractionCounter",{"@type":66},"ViewAction",{"@type":68,"mainEntity":69},"FAQPage",[70,76,80],{"name":71,"@type":72,"acceptedAnswer":73},"Why is synthetic speech harder to distinguish using only listening?","Question",{"text":74,"@type":75},"The paper notes that the public often believes synthetic speech is difficult to detect by listening alone, while only a smaller fraction knows it can be clearly distinguished using appropriate methods.","Answer",{"name":77,"@type":72,"acceptedAnswer":78},"What property of Japanese synthetic speech leads to the proposed detection advantage?",{"text":79,"@type":75},"Because synthetic speech is learned from a limited set of training spectra, the variety of vowel spectra is intrinsically limited compared with natural speech.",{"name":81,"@type":72,"acceptedAnswer":82},"How does the method separate synthetic from natural speech?",{"text":83,"@type":75},"It extracts and normalizes vowel mora spectra from the same read text, treats them as probability density functions, computes Wasserstein distances, and applies persistent homology topological mapping to decompose spectra into distinct clusters for synthetic versus natural speech.","https://schema.org",{"og:url":50,"og:type":86,"og:title":13,"og:site_name":57,"og:description":14},"article",{"robots":88,"canonical":50},"index,follow",{"doc_id":7,"site_id":23},{"code":4,"msg":5,"data":91},[92,96,100,104,109,114,119,122,126,129,133],{"id":20,"doc_module":4,"doc_module_name":44,"category_name":93,"show_sort_weight":94,"slug":95},"Story & Novel",90,"story-novel",{"id":45,"doc_module":4,"doc_module_name":44,"category_name":97,"show_sort_weight":98,"slug":99},"Literature",80,"literature",{"id":51,"doc_module":4,"doc_module_name":44,"category_name":101,"show_sort_weight":102,"slug":103},"Exam",70,"exam",{"id":105,"doc_module":4,"doc_module_name":44,"category_name":106,"show_sort_weight":107,"slug":108},5,"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":44,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":44,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":44,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":44,"category_name":124,"show_sort_weight":27,"slug":125},9,"Religion & Spirituality","religion-spirituality",{"id":27,"doc_module":4,"doc_module_name":44,"category_name":127,"show_sort_weight":27,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":44,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":44,"category_name":135,"show_sort_weight":105,"slug":136},19,"General","general"]