[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83859-en":3,"doc-seo-83859-105":30,"detail-sidebar-cat-0-en-105":84},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83859,8796095462418,"Noah","https://ap-avatar.wpscdn.com/avatar/80000253c1241d02b47?x-image-process=image/resize,m_fixed,w_180,h_180&k=1778826106357471780",8,"Research & Report","SynSFX Multi-Model Sound Effects Synthesis Dataset for Deepfake Detection and Evaluation","Audio deepfake detection has advanced, yet existing detectors generalize poorly to synthetic sound effects. SynSFX addresses this gap by providing a large-scale corpus of 43,374 clips, including 26,452 synthetic and 16,922 real recordings. The dataset spans seven popular text-to-audio models and emphasizes non-speech forensics for environmental sounds, Foley, and ambient textures, enabling controlled evaluation beyond speech-focused benchmarks.","SynSFX: Multi-Model Sound Effects Synthesis Dataset for Deepfake Detection and Evaluation  \nLinxi Li∗†* , Yuncong Yu†* , Qianwei Guo†, Liwei Jin†, Yechen Wang†, Carsten Maple∗  \n∗ University of Warwick, Coventry, United Kingdom  \n†OfSpectrum, Inc., Los Angeles, CA, USA  \narXiv :2607 .04848v 1 [ cs . SD] 6 Jul 2026  \nAbstract—While audio deepfake detection has advanced significantly, representative detectors show limited generalization to synthetic sound effects. Existing environmental audio datasets such as EnvSDD provide important initial resources, but remain limited in scale and generation provenance for studying isolated sound-effect deepfakes. To support this direction, we present SynSFX, a large-scale corpus of 43374 clips (26452 synthetic, 16922 real) spanning 7 popular text-to-audio models.  \nIndex Terms—audio deepfake detection, sound effects, spoofing, non-speech audio, dataset  \nI. INTRODUCTION  \nRapid advancements in deep generative modeling for audio—encompassing both speech and sound effects—have significantly outpaced the development of robust deepfake detection countermeasures [1] . State-of-the-art text-to-audio frameworks now synthesize high-fidelity acoustic events that are frequently perceptually indistinguishable from genuine recordings [2] . However, contemporary deepfake detection literature remains disproportionately anchored to speech synthesis and voice cloning paradigms [3] . Consequently, a critical research gap persists regarding non-speech audio forensics, specifically encompassing environmental sounds, Foley, and ambient textures. Despite being equally susceptible to malicious manipulation as artificial speech, the targeted detection of synthetic sound effects remains a fundamentally underexplored domain.  \nHowever, sound effects differ fundamentally from speech [4] . Unlike speech, which contains lexical and prosodic cues that detectors can exploit, non-speech audio such as sound effects and environmental sounds lacks structured linguistic information. Detection therefore relies purely on acoustic characteristics which modern generative models can closely imitate. As a result, systems trained primarily on speech data often fail to generalize to non-speech scenarios.  \nThis limitation is consequential. Sound effects are widely used in gaming, media production, live streaming, and safetycritical systems [5] . As generative models improve, the risk of misuse increases—from subtle manipulation in entertainment to injection of deceptive environmental sounds. To address this gap, we introduce SynSFX (Synthetic Sound Effects), a largescale corpus for non-speech audio deepfake detection. The dataset includes diverse sound effects and environmental audio  \n*Equal contribution.  \ngenerated by multiple modern models, establishing a comprehensive benchmark beyond speech-focused evaluations.  \nThe primary contributions of this work are summarized as follows:  \n• A Large-Scale Benchmark with Diagnostic Control: We release SynSFX, a comprehensive 178-hour corpus encompassing seven diverse text-to-audio architectures. Beyond providing massive scale and acoustic diversity, SynSFX uniquely features a specialized Shared Prompt Subset. The shared-prompt subset provides a controlled resource for future prompt-matched analyses of generator-dependent artifacts.  \n• Identification of Catastrophic Forgetting: We empirically reveal that adapting speech-centric detectors exclusively to sound effects causes a catastrophic degradation in speech spoofing detection. We demonstrate that while joint-domain training mitigates this forgetting, critical generalization bottlenecks remain.  \n• Exposing the Illusion of Generalization: Through rigorous zero-shot evaluations on unseen generation architectures and feature space visualizations (t-SNE), we show that current models overfit to known synthesis artifacts rather than learning universal acoustic anomalies, establishing a baseline for future generalized audio forensics.  \nII. RE","cbCaiaUKwmC9dXGv","https://ap.wps.com/l/cbCaiaUKwmC9dXGv","pdf",270007,5,1,7,"English","en",105,"# Introduction\n## Sound effects vs. speech detection gap\n## SynSFX dataset contributions\n# Related Work\n## Speech deepfake detection\n## Advancements in audio generation models\n## Deepfake detection for general audio and sound effects","[{\"question\":\"What findings does the paper report when adapting speech-centric detectors to sound effects?\",\"answer\":\"Adapting speech-focused detectors exclusively to sound effects causes catastrophic degradation in speech spoofing detection; joint-domain training mitigates forgetting but leaves generalization bottlenecks.\"}]",1784191027,18,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":79,"head_meta":81,"extra_data":83,"updated_unix":28},"synsfx-multi-model-sound-effects-synthesis-dataset-for-deepfake-detection-and-evaluation","",{"@graph":36,"@context":78},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/synsfx-multi-model-sound-effects-synthesis-dataset-for-deepfake-detection-and-evaluation/83859/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72],{"name":73,"@type":74,"acceptedAnswer":75},"What findings does the paper report when adapting speech-centric detectors to sound effects?","Question",{"text":76,"@type":77},"Adapting speech-focused detectors exclusively to sound effects causes catastrophic degradation in speech spoofing detection; joint-domain training mitigates forgetting but leaves generalization bottlenecks.","Answer","https://schema.org",{"og:url":52,"og:type":80,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":82,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":85},[86,90,94,98,102,107,111,114,119,122,126],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":87,"show_sort_weight":88,"slug":89},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":91,"show_sort_weight":92,"slug":93},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Comic",60,"comic",{"id":103,"doc_module":4,"doc_module_name":46,"category_name":104,"show_sort_weight":105,"slug":106},6,"Technology",50,"technology",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":112,"slug":113},30,"research-report",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},9,"Religion & Spirituality",20,"religion-spirituality",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":120,"show_sort_weight":117,"slug":121},"World Cup","world-cup",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":123,"slug":125},10,"Lifestyle","lifestyle",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":20,"slug":129},19,"General","general"]