[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85102-en":3,"doc-seo-85102-105":28,"detail-sidebar-cat-0-en-105":90},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":11,"language":21,"language_code":22,"site_id":23,"html_lang":22,"table_of_contents":24,"faqs":25,"seo_title":13,"seo_description":14,"update_tm":26,"read_time":27},85102,687197207057,"Sage","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","Why Do You Say It Like That? A Phoneme-Level Framework for Explainable Speech Deepfake Detection","As speech deepfake detection accuracy improves using self-supervised representations such as wav2vec 2.0 and HuBERT, the rationale behind bona fide versus deepfake classifications remains unclear. A phoneme-level analysis framework is introduced to connect model outputs to measurable phonetic units via a post-hoc Grad-CAM explainability pipeline with speech recognition alignment. Experiments on ASVspoof 5 match comparable detection performance while yielding statistically significant, attack- and speaker-dependent phonetic cues interpretable by humans across speakers and spoofing conditions.","Why Do You Say It Like That? A Phoneme-Level Framework for Explainable Speech Deepfake Detection  \nAnna Taylor1 , Michele Panariello 1 , Massimiliano Todisco 1 , Chiara Galdi 1 , Nicholas Evans 1 , Driss Matrouf2  \n1EURECOM, Sophia Antipolis, France  \n2Laboratoire Informatique d’Avignon, Avignon Universit, France  \n[1](1 firstname.lastname@eurecom.fr)[ firstname.lastname@eurecom.fr](1 firstname.lastname@eurecom.fr)  \n[2](2 driss.matrouf@univ-avignon.fr)[ driss.matrouf@univ-avignon.fr](2 driss.matrouf@univ-avignon.fr)  \narXiv :2607 .08586v1 [ ee ss .AS] 9 Jul 2026  \nAbstract—As the accuracy of speech deepfake detection improves with the use of self-supervised representations such as wav2vec 2.0 and HuBERT, understanding why the speech is classified as bona fide or deepfake remains an open challenge. In pursuit of more trustworthy and interpretable artificial intelligence, we introduce a phoneme-level analysis framework that connects model predictions to measurable phonetic units. Our post-hoc explainability method is generally applicable toa variety of speech deepfake detection systems based on convolutional neural networks since it leverages Gradient-weighted Class Activation Mapping in conjunction with speech recognition to generate saliency maps aligned with phonemes and pauses. This pipeline reveals statistically significant attack-and speakerdependent phonetic cues associated with spoofed speech in terms that humans can understand. Experiments using ASVspoof 5 show comparable detection performance to similar architectures while providing linguistic interpretations across speakers and spoofing conditions.  \nIndex Terms—deepfake detection, speech processing, explainable artificial intelligence  \nI. INTRODUCTION  \nAs progress in the field of generative speech technology has pushed the deceptive capabilities of voice conversion and text-to-speech to greater levels, it is consequently evermore important to distinguish between bona fide and spoofed speech. Modern speech deepfake detection systems manage to achieve strong performance by leveraging self-supervised (SSL) speech representations extracted using large foundation models such as wav2vec 2.0 [1], HuBERT [2], and WavLM [3] . Despite these advances, the predictions of such systems remain largely difficult to understand with humanlike reasoning [4] and they lack human-perceptible cues that would provide natural explanations for bona fide or spoof classifications. Even though a deepfake detector may successfully classify an utterance, it is often unclear which characteristics of the speech signal contributed to that prediction.  \nThis work was supported by the COMPROMIS project (ANR22-PECY- 0011) funded by a French government grant managed by the Agence Nationale de la Recherche under the France 2030 program. This work has been submitted to the IEEE for possible publication. Copyright may be transferred without notice, after which this version may no longer be accessible.  \nThis lack of interpretability presents a challenge for the trustworthiness and transparency of speech deepfake detection systems. In applications where security and fairness are critical, understanding why a system reached a particular classification is crucial not only for faith in the decision itself, but also for future adjustment of the system. However, explanations based solely on low-level acoustic representations are often difficult to interpret, particularly for speech signals where the structure is more naturally described in linguistic terms, such as phonemes.  \nPhonemes provide a useful intermediate representation between raw acoustics and human interpretation of spoken language. In accordance with the International Phonetic Association and leading scholars, the definition of a phoneme is strictly limited to the assigned categorical representation of an uttered sound in a given language which does not incorporate its acoustic realization or varied pronunciation across dialects and spontaneous s","cbCaiuMjmyz9l2AD","https://ap.wps.com/l/cbCaiuMjmyz9l2AD","pdf",1099225,1,"English","en",105,"# Abstract\n# Introduction\n## Motivation and interpretability challenge\n## Phonemes as an intermediate representation\n## Prior work and open questions\n## Proposed phoneme-level explainability framework\n# Method Overview\n## Self-supervised front-end and CNN classifier\n## Post-hoc Grad-CAM with forced alignment\n# Experiments and Results","[{\"question\":\"Why is interpretability still difficult in modern speech deepfake detection systems?\",\"answer\":\"Even when detectors classify speech accurately, it is often unclear which characteristics of the signal drive the prediction, and explanations based only on low-level acoustic representations are hard to interpret for humans.\"},{\"question\":\"What does the proposed framework explain at a phoneme level?\",\"answer\":\"It aligns temporal saliency maps with phoneme boundaries obtained via automatic speech recognition and forced alignment, enabling analysis of which phonetic units contribute to bona fide or deepfake decisions.\"},{\"question\":\"How is the framework validated and what outcomes are reported on ASVspoof 5?\",\"answer\":\"Experiments on ASVspoof 5 show detection performance comparable to similar architectures, while providing linguistic interpretations with statistically significant attack- and speaker-dependent phonetic cues across conditions.\"}]",1784201113,20,{"code":4,"msg":29,"data":30},"ok",{"site_id":23,"language":22,"slug":31,"title":13,"keywords":32,"description":14,"schema_data":33,"social_meta":85,"head_meta":87,"extra_data":89,"updated_unix":26},"why-do-you-say-it-like-that-a-phoneme-level-framework-for-explainable-speech-deepfake-detection","",{"@graph":34,"@context":84},[35,52,67],{"@type":36,"itemListElement":37},"BreadcrumbList",[38,42,46,49],{"item":39,"name":40,"@type":41,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":43,"name":44,"@type":41,"position":45},"https://docshare.wps.com/document/","Document",2,{"item":47,"name":12,"@type":41,"position":48},"https://docshare.wps.com/document/research-report/",3,{"item":50,"name":13,"@type":41,"position":51},"https://docshare.wps.com/document/why-do-you-say-it-like-that-a-phoneme-level-framework-for-explainable-speech-deepfake-detection/85102/",4,{"url":50,"name":13,"@type":53,"author":54,"headline":13,"publisher":56,"fileFormat":59,"inLanguage":22,"description":14,"dateModified":60,"datePublished":61,"encodingFormat":59,"isAccessibleForFree":62,"interactionStatistic":63},"DigitalDocument",{"name":9,"@type":55},"Person",{"url":39,"name":57,"@type":58},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":64,"interactionType":65,"userInteractionCount":20},"InteractionCounter",{"@type":66},"ViewAction",{"@type":68,"mainEntity":69},"FAQPage",[70,76,80],{"name":71,"@type":72,"acceptedAnswer":73},"Why is interpretability still difficult in modern speech deepfake detection systems?","Question",{"text":74,"@type":75},"Even when detectors classify speech accurately, it is often unclear which characteristics of the signal drive the prediction, and explanations based only on low-level acoustic representations are hard to interpret for humans.","Answer",{"name":77,"@type":72,"acceptedAnswer":78},"What does the proposed framework explain at a phoneme level?",{"text":79,"@type":75},"It aligns temporal saliency maps with phoneme boundaries obtained via automatic speech recognition and forced alignment, enabling analysis of which phonetic units contribute to bona fide or deepfake decisions.",{"name":81,"@type":72,"acceptedAnswer":82},"How is the framework validated and what outcomes are reported on ASVspoof 5?",{"text":83,"@type":75},"Experiments on ASVspoof 5 show detection performance comparable to similar architectures, while providing linguistic interpretations with statistically significant attack- and speaker-dependent phonetic cues across conditions.","https://schema.org",{"og:url":50,"og:type":86,"og:title":13,"og:site_name":57,"og:description":14},"article",{"robots":88,"canonical":50},"index,follow",{"doc_id":7,"site_id":23},{"code":4,"msg":5,"data":91},[92,96,100,104,109,114,119,122,126,129,133],{"id":20,"doc_module":4,"doc_module_name":44,"category_name":93,"show_sort_weight":94,"slug":95},"Story & Novel",90,"story-novel",{"id":45,"doc_module":4,"doc_module_name":44,"category_name":97,"show_sort_weight":98,"slug":99},"Literature",80,"literature",{"id":51,"doc_module":4,"doc_module_name":44,"category_name":101,"show_sort_weight":102,"slug":103},"Exam",70,"exam",{"id":105,"doc_module":4,"doc_module_name":44,"category_name":106,"show_sort_weight":107,"slug":108},5,"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":44,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":44,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":44,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":44,"category_name":124,"show_sort_weight":27,"slug":125},9,"Religion & Spirituality","religion-spirituality",{"id":27,"doc_module":4,"doc_module_name":44,"category_name":127,"show_sort_weight":27,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":44,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":44,"category_name":135,"show_sort_weight":105,"slug":136},19,"General","general"]