[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83717-en":3,"doc-seo-83717-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83717,4398048949847,"Eliana","https://ap-avatar.wpscdn.com/avatar/400002536579ef2da7f?_k=1778318612642679267",8,"Research & Report","DETECT-3B-Omni is Agnostic of Content and Demographics","A trustworthy, GDPR-compliant deepfake audio detector must decide from acoustic artifacts, not from the spoken message or the speaker’s identity. DETECT-3B-Omni is evaluated in a large-scale semantic-independence study using 10,240 audio samples from diverse U.S. English speakers across 30 states, generated by 8 AI voice-cloning systems. Equivalence testing shows detection accuracy differences across content type and speaker demographics are at most 2 percentage points at 99% confidence.","arXiv :2607 .034 18v 1 [ cs . SD] 3 Jul 2026  \nDETECT-3B-Omni is Agnostic of Content and Demographics  \nNicolas M. Müller 1 , Aditya Tirumala Bukkapatnam 1 , Dominik Schnieders2 , Zohaib Ahmed 1  \n1 Resemble AI, Mountain View, CA, USA {nicolas,aditya,zohaib}@resemble.ai  \n2 Deutsche Telekom, Bonn, Germany Dominik .Schnieders@telekom .de  \nAbstract  \nA trustworthy and GDPR-compliant deepfake audio detector must base its decisionson acoustic artifacts, not on what is being said or who is speaking. We present a largescale study of semantic independence for Resemble AI’s detector, DETECT-3B-Omni. Using 10,240 audio samples from diverse US English speakers across 30 states, generated through 8 different AI voice-cloning systems, we test whether detection accuracy depends on spoken content (benign versus malicious), speaker gender, speaker age, or speaker region. Using equivalence testing, our results show that the accuracy difference between any two of these groups is at most 2 percentage points, at 99% confidence. The detector therefore identifies AI-generated audio with equivalent accuracy regardless of what the audio says or who the speaker is.  \n1 Introduction  \nAudio deepfake detection systems are increasingly deployed to protect phone calls and voice authentication against AI-generated fraud. A key use case is real-time monitoring of telephone calls to flag AI-generated speech before it can cause harm. Under the EU General Data Protection Regulation (GDPR) and applicable telecommunication law, deploying such a system requires demonstrating that it processes only the information strictly necessary for its purpose: detecting whether audio is AI-generated. In particular, the detector must not make decisions based on the content of the speech (what is said) or the identity of the speaker (who is speaking), as this would constitute unnecessary processing of personal data.  \nA detector that is semantically independent satisfies this requirement: its output depends only on acoustic artifacts that distinguish human speech from AI-generated speech, not on the words being spoken or the characteristics of the speaker. If a detector’s accuracy changed depending on the topic of conversation (for example, performing differently on financial discussions versus casual speech), it would effectively be “listening to” the content, raising both fairness and GDPR compliance concerns. Similarly, if detection accuracy varied across speaker demographics (gender, age, or regional accent), the system would be unreliable for certain populations and potentially discriminatory.  \nWe present a controlled, large-scale study to verify that DETECT-3B-Omni exhibits no such dependence. We construct a dataset of 10,240 audio samples covering:  \n• 640 unique sentences, split evenly between benign (everyday business communica-  \ntion) and malicious (social engineering) content,  \n• recordings by diverse native US English speakers, balanced across gender and spanning ages 20–55 and 30 US states,  \n• fake audio generated by 8 state-of-the-art open-source text-to-speech voice-cloning models.  \nWe classify every sample using the DETECT-3B-Omni API and compare detection accuracy across content categories and demographic groups. The goal is straightforward: if the detector is semantically independent, accuracy should be the same regardless of how we split the data. This provides evidence that the system is suitable for deployment in GDPR-regulated environments, where it can monitor calls for deepfake audio without processing the semantic content of conversations.  \n2 Data  \n2.1 Sentence Design  \nWe designed 640 English sentences in two categories of 320 each:  \n• Benign: standard business and technical communication (e.g., invoice confirmations, meeting scheduling, account inquiries) .  \n• Malicious: social engineering and fraud scenarios (e.g. , urgent payment demands, credential phishing, authority impersonation) .  \nSentences are of comparable length (115–150 charac","cbCaiunz5wqDXK6V","https://ap.wps.com/l/cbCaiunz5wqDXK6V","pdf",518954,5,1,9,"English","en",105,"# Abstract\n# Introduction\n# Data\n## Sentence Design\n## Real Audio\n## Fake Audio\n# Experiment Design","[{\"question\":\"What does DETECT-3B-Omni aim to be independent of?\",\"answer\":\"DETECT-3B-Omni is designed so its decisions rely on acoustic artifacts rather than on what is being said or who is speaking.\"},{\"question\":\"How was the study dataset constructed?\",\"answer\":\"The dataset includes 10,240 samples: 640 English sentences recorded by 8 native U.S. speakers and synthesized into spoofed audio using 8 open-source voice-cloning/TTS models, yielding matching real and fake counts.\"},{\"question\":\"What do the equivalence testing results indicate?\",\"answer\":\"Accuracy differences between groups defined by content category and speaker demographics are at most 2 percentage points with 99% confidence, supporting semantic independence.\"}]",1784189947,23,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"detect-3b-omni-is-agnostic-of-content-and-demographics","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/detect-3b-omni-is-agnostic-of-content-and-demographics/83717/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What does DETECT-3B-Omni aim to be independent of?","Question",{"text":76,"@type":77},"DETECT-3B-Omni is designed so its decisions rely on acoustic artifacts rather than on what is being said or who is speaking.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How was the study dataset constructed?",{"text":81,"@type":77},"The dataset includes 10,240 samples: 640 English sentences recorded by 8 native U.S. speakers and synthesized into spoofed audio using 8 open-source voice-cloning/TTS models, yielding matching real and fake counts.",{"name":83,"@type":74,"acceptedAnswer":84},"What do the equivalence testing results indicate?",{"text":85,"@type":77},"Accuracy differences between groups defined by content category and speaker demographics are at most 2 percentage points with 99% confidence, supporting semantic independence.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,110,115,120,123,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":20,"slug":137},19,"General","general"]