[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86306-en":3,"doc-seo-86306-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86306,13056703020460,"Valentina","https://ap-avatar.wpscdn.com/avatar/be000253dac470eee5d?_k=1778207105932848923",8,"Research & Report","VoxENES 2026 Benchmarking Generalization of Speech Spoofing Detectors Against LLM-Era TTS and Voice Conversion","VoxENES 2026 targets a growing evaluation mismatch between legacy speech-spoofing benchmarks and modern LLM-driven text-to-speech (TTS) and voice conversion (VC) generators. The dataset models temporal generalization gap by supplying 53,628 bilingual (English/Spanish) audio samples created with 10 contemporary synthesis methods and tested under 10 standardized postprocessing conditions. Eight pretrained detectors are benchmarked without fine-tuning, revealing substantial out-of-distribution performance collapse, with the best model reaching only 28.98% EER overall, indicating reliance on brittle artifacts and motivating robust countermeasure development.","VoxENES 2026: Benchmarking Generalization of Speech Spoofing Detectors Against LLM-Era TTS and Voice Conversion  \nAastha Sharma, Guangjing Wang  \nUniversity of South Florida, Tampa, FL, USA  \n[aasthasharma@usf.edu](aasthasharma@usf.edu) , [guangjingwang@usf.edu](guangjingwang@usf.edu)  \narXiv :2607 . 1 1706v 1 [ cs . SD] 13 Jul 2026  \nAbstract  \nModern LLM-driven text-to-speech (TTS) and voice conversion (VC) systems produce synthetic speech that differs from the generators represented in many legacy spoofing benchmarks. This mismatch creates a temporal generalization gap that can overestimate detector robustness under real-world postprocessing conditions. We bridge this gap by introducing VoxENES 2026, a bilingual (English and Spanish) benchmark of 53,628 audio samples generated using 10 contemporary speech synthesis methods and evaluated under 10 standardized postprocessing conditions. Using VoxENES 2026, we benchmark eight pretrained detectors without fine-tuning and observe substantial performance degradation: the best model achieves 28.98% EER overall, while most perform near or below random chance across modern generators and perturbations. Our results highlight the reliance on brittle artifacts in current detectors and establish VoxENES 2026 as a practical testbed for developing robust audio spoofing countermeasures.  \nIndex Terms: audio deepfake detection, speech spoofing detection, benchmark dataset  \n1. Introduction  \nRobust speech spoofing and deepfake detection are essential to preserve trust in speech-based authentication [1, 2, 3, 4] . This need is growing as voice becomes a biometric and a control channel for speech-driven agents and assistive technologies. If synthetic speech becomes indistinguishable from genuine speech, detection failures not only compromise security protocols but also erode trust in audio evidence.  \nMany benchmark datasets are proposed for speech spoofing and deepfake detection evaluation. For example, the ASVspoof challenge series has been the primary driver of spoofing countermeasure development. ASVspoof 2019 [5] introduces logical access (LA) with TTS and VC, and physical access (PA) with replay tracks. ASVspoof 2021 [6] adds the deepfake task targeting compressed manipulated speech, and ASVspoof 5 [7] introduces crowdsourced data with adversarial attacks at scale. In addition to ASVspoof, WaveFake [8] provides a multilingual dataset from six neural vocoder architectures. The In-the-Wild dataset [9] includes real-world deepfakes of celebrities. The MLAAD [10] dataset expands coverage to 23 languages and 54 TTS models. The VoiceWukong [11] benchmarks 12 detectors against 34 commercial and open-source tools with postprocessing manipulations.  \nYet, existing benchmarks primarily rely on speech synthesis systems before 2024 and fail to capture artifact patterns produced by the modern large language model (LLM)-driven generation pipelines. For example, text-to-speech (TTS) designs include autoregressive language-model-based synthesis, such  \nas VoxCPM [12] and Qwen3-TTS [13]; flow-matching models, including GLM-TTS [14], CosyVoice 3 [15], and Chatterbox [16]; diffusion-based systems like FlashLabs Chroma [17]; and hybrid DiT architectures, such as VibeVoice [18] . For voice conversion (VC), zero-shot approaches such as Seed-VC [19], tone-color extraction methods such as OpenVoice v2 [20], and retrieval-based systems like RVC v2 [21] have substantially improved naturalness and speaker similarity.  \nEvaluation on temporally stale benchmarks can overestimate real-world robustness as speech synthesis techniques improve. LLM-driven TTS and VC systems produce synthetic audio with acoustic characteristics that differ substantially from earlier spoofing corpora, creating a data drifting issue for existing detectors. Deepfake detection models that perform well on older benchmarks may fail when deployed against newer deepfake generators. This mismatch reflects a cat-and-mouse dynamic in which detector","cbCaijrPvHw4WXoE","https://ap.wps.com/l/cbCaijrPvHw4WXoE","pdf",1358551,10,1,5,"English","en",105,"# Introduction\n## Related Work and Existing Benchmarks\n## Problem: Temporal Generalization Gap and Data Drift\n## VoxENES 2026 Dataset and Postprocessing Setup\n## Out-of-Distribution Evaluation Results","[{\"question\":\"What problem does VoxENES 2026 address in speech spoofing detection?\",\"answer\":\"It addresses temporal generalization gap: detectors benchmarked on pre-2024 synthesis struggle against LLM-era TTS/VC because the generated acoustic characteristics and postprocessing differ from legacy data.\"},{\"question\":\"How is VoxENES 2026 constructed?\",\"answer\":\"VoxENES 2026 provides 53,628 English and Spanish audio samples generated using 10 contemporary speech synthesis methods (7 TTS, 3 VC) and evaluated under 10 standardized postprocessing conditions.\"},{\"question\":\"What do the benchmark results show about existing detectors?\",\"answer\":\"Without fine-tuning, eight pretrained detectors show substantial degradation; the best achieves 28.98% EER overall, while most perform near or below random chance across modern generators and perturbations.\"}]",1784210350,13,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"voxenes-2026-benchmarking-generalization-of-speech-spoofing-detectors-against-llm-era-tts-and-voice-conversion","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/voxenes-2026-benchmarking-generalization-of-speech-spoofing-detectors-against-llm-era-tts-and-voice-conversion/86306/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-26","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What problem does VoxENES 2026 address in speech spoofing detection?","Question",{"text":76,"@type":77},"It addresses temporal generalization gap: detectors benchmarked on pre-2024 synthesis struggle against LLM-era TTS/VC because the generated acoustic characteristics and postprocessing differ from legacy data.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How is VoxENES 2026 constructed?",{"text":81,"@type":77},"VoxENES 2026 provides 53,628 English and Spanish audio samples generated using 10 contemporary speech synthesis methods (7 TTS, 3 VC) and evaluated under 10 standardized postprocessing conditions.",{"name":83,"@type":74,"acceptedAnswer":84},"What do the benchmark results show about existing detectors?",{"text":85,"@type":77},"Without fine-tuning, eight pretrained detectors show substantial degradation; the best achieves 28.98% EER overall, while most perform near or below random chance across modern generators and perturbations.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,110,115,120,123,128,131,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":20,"slug":133},"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":22,"slug":137},19,"General","general"]