[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84470-en":3,"doc-seo-84470-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},84470,687197100911,"Himbo","https://ap-avatar.wpscdn.com/avatar/a000239b6f1da00475?x-image-process=image/resize,m_fixed,w_180,h_180&k=1782698725881665579",8,"Research & Report","Towards Robust Speech Deepfake Detection via Human-Inspired Reasoning","Modern generative audio systems can be misused to impersonate individuals and obtain private information, driving the evolution of speech deepfake detection (SDD). Existing SDD approaches often generalize poorly to new audio domains and generators, and they lack interpretability with human-like reasoning that explains whether speech is bona fide or spoof. HIR-SDD addresses this by combining large audio language models with chain-of-thought reasoning trained via a human-annotated dataset, improving detection effectiveness and prediction justification.","Towards Robust Speech Deepfake Detection via Human-Inspired Reasoning  \nArtem Dvirniak 1, Evgeny Kushnir2 ,3 ,4, Dmitrii Tarasov3 ,5, Artem Iudin6, Oleg Kiriukhin7, Mikhail  \nPautov2 ,8, Dmitrii Korzh2 ,4 ,6 ,∗∗, Oleg Y. Rogov2 ,4 ,6  \n1MIRAI, 2AXXX, 3HSE, 4Applied AI Institute, 5Fusion Brain Lab, AXXX,  \n6MTUCI, 7 City University of Hong Kong, 8Trusted AI Research Center, RAS  \n[d.s.korzh@mtuci.ru](d.s.korzh@mtuci.ru)  \narXiv :2603 . 10725v 3 [ cs . SD] 13 Jul 2026  \nAbstract  \nThe modern generative audio models can be used by an adversary in an unlawful manner, specifically, to impersonate other people to gain access to private information. To mitigate this issue, speech deepfake detection (SDD) methods started to evolve. Unfortunately, current SDD methods generally suffer from the lack of generalization to new audio domains and generators. More than that, they lack interpretability, especially human-like reasoning that would naturally explain the attribution of a given audio to the bona fide or spoof class and provide human-perceptible cues. In this paper, we propose HIR-SDD, a novel SDD framework that combines the strengths of Large Audio Language Models (LALMs) with the chain-of-thought reasoning derived from the novel proposed human-annotated dataset. Experimental evaluation demonstrates both the effectiveness of the proposed method and its ability to provide reasonable justifications for predictions.  \nIndex Terms: deepfake detection, voice anti-spoofing, audio LLM, reasoning, benchmark  \n1. Introduction  \nContests, such as ASVspoof [1, 2], ADD [3, 4], and Singfake [5, 6] drive the progress in audio and speech deepfake detection (SDD) research, providing high-quality deepfake data and fair evaluation protocols. SDD and spoofing-aware speaker verification (SASV) research primarily focuses on the architecture design, including task-specific front-ends [7], graph-attention networks [8], self-supervised (SSL) audio encoders [9, 10], and architecture modifications [11, 12] to improve the empirical performance of the models. Additionally, augmentation strategies [13], optimizer choices [14], and representation learning-based approaches [15, 16, 17] are also investigated in SDD research.  \nUnfortunately, SDD remains challenging due to distribution shifts across spoofing methods, speech domains, and transformations, as evidenced by results from the aforementioned contests. A common strategy to improve empirical performance is to increase training diversity and model capacity. However, SDD systems still fail to generalize to unseen domains, highlighting the need for more robust and explainable detection approaches. Such tools are particularly important for risksensitive applications, such as biometrics and banking, yet this direction remains underexplored in SDD research.  \nRecently, Large Language Models (LLMs) [18, 19] and Large Audio Language Models (LALMs), such as SALMONN [20], Qwen2-Audio [21], and AudioFlamingo 3 [22], have demonstrated strong reasoning capabilities. Chain-of-thought (CoT) [23] and related reason-  \n**indicates the corresponding author.  \ning methods [19] often improve performance by extracting intermediate rationales. Although CoT mainly acts as an empirical mechanism and does not necessarily reflect human reasoning, careful training and grounding can produce explanations that are consistent with the input and useful for analysis. Similar ideas have been explored for audio question answering and captioning (including music understanding) [24, 25], disease detection from speech [26], emotion recognition [27], and interpretable audio quality assessment and hard-label SDD [28] .  \nHowever, reasoning-based approaches remain relatively uncommon in SDD. One practical limitation is the lack of opensource datasets with high-quality human explanations for training and evaluation. Moreover, existing LALMs, even state-ofthe-art, drastically underperform in zero- and few-shot SDD, thereby limiting the reliability of","cbCaiv4hwJ6ZyJF5","https://ap.wps.com/l/cbCaiv4hwJ6ZyJF5","pdf",220433,1,6,"English","en",105,"# Introduction\n# Related work","[{\"question\":\"What problem does this paper address in speech deepfake detection?\",\"answer\":\"It targets limited generalization of current SDD methods to new domains and generators, as well as insufficient interpretability, especially the lack of human-like reasoning explaining bona fide vs spoof attribution.\"},{\"question\":\"What is HIR-SDD and how does it improve detection?\",\"answer\":\"HIR-SDD combines large audio language models with chain-of-thought reasoning. It uses a human-annotated dataset for CoT training and applies hard-label and CoT-supervised fine-tuning, grounding, and reinforcement learning.\"},{\"question\":\"What resources and results does the paper report?\",\"answer\":\"The work introduces a human-annotated reasoning dataset for 41k bona fide and spoof samples and proposes hard-label plus CoT pipelines for strong countermeasure performance with reasoning explainability, along with improvement and evaluation strategies for reasoning-capable SDD models.\"}]",1784195852,15,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"towards-robust-speech-deepfake-detection-via-human-inspired-reasoning","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/towards-robust-speech-deepfake-detection-via-human-inspired-reasoning/84470/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-21","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does this paper address in speech deepfake detection?","Question",{"text":75,"@type":76},"It targets limited generalization of current SDD methods to new domains and generators, as well as insufficient interpretability, especially the lack of human-like reasoning explaining bona fide vs spoof attribution.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What is HIR-SDD and how does it improve detection?",{"text":80,"@type":76},"HIR-SDD combines large audio language models with chain-of-thought reasoning. It uses a human-annotated dataset for CoT training and applies hard-label and CoT-supervised fine-tuning, grounding, and reinforcement learning.",{"name":82,"@type":73,"acceptedAnswer":83},"What resources and results does the paper report?",{"text":84,"@type":76},"The work introduces a human-annotated reasoning dataset for 41k bona fide and spoof samples and proposes hard-label plus CoT pipelines for strong countermeasure performance with reasoning explainability, along with improvement and evaluation strategies for reasoning-capable SDD models.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,114,119,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":21,"doc_module":4,"doc_module_name":45,"category_name":111,"show_sort_weight":112,"slug":113},"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":45,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":45,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":45,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]