[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84486-en":3,"doc-seo-84486-105":29,"detail-sidebar-cat-0-en-105":90},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},84486,549758146520,"Patrick","https://ap-avatar.wpscdn.com/avatar/80002397d8c0411e94?_k=1775819394049821470",8,"Research & Report","Are DeepFakes Realistic Enough? Exploring Semantic Mismatch as a Novel Challenge","DeepFake detection often uses binary or four-class audio-visual assumptions, overlooking variability in manipulation sources and, critically, semantic consistency across modalities. The proposed evaluation extends the four-class framework by introducing Real Audio–Real Video with Semantic Mismatch (RARV-SMM), where both streams are authentic but originate from different events, contexts, or narratives. Experiments with FakeAVCeleb test state-of-the-art robustness, reveal limitations, and introduce RARV-SMM variants to expose architectural vulnerabilities. A semantic reinforcement strategy with frozen ImageBind embeddings improves analysis of semantic coherence for realistic detection scenarios.","Are DeepFakes Realistic Enough? Exploring Semantic Mismatch as a Novel  \nChallenge  \nSharayu Nilesh Deshmukh, Kailash A. Hambarde, Joana C. Costa, Hugo Proenc¸a, and Tiago Roxo Instituto de Telecomunicac¸ es, Universidade da Beira Interior, Portugal  \n{d.sharyu.nilesh, kailash.hambarde, joana.cabral.costa, hugomcp, [tiago.roxo](tiago.roxo}@ubi.pt)[}](tiago.roxo}@ubi.pt)[@ubi.pt](tiago.roxo}@ubi.pt)  \narXiv :2604 .28022v2 [ cs .CV] 13 Jul 2026  \nAbstract  \nCurrent DeepFake detection scenarios are mostly binary, yet data manipulation can vary across audio, video, or both, whose variability is not captured in binary settings. Fourclass audio-visual formulations address this by discriminating manipulation type, but introduce an unresolved problem: models may rely solely on data source integrity to detect DeepFakes without evaluating their semantic consistency. If the DeepFake origin is not in the data source but in its content, can semantic mismatch be assessed by the stateof-the-art? This paper proposes a new evaluation setup, extending the four-class formulation by explicitly modeling semantic-level inconsistency between authentic modalities with the introduction of a new class: Real Audio –Real Video with Semantic Mismatch (RARV-SMM). We assess the robustness of state-of-the-art models in this new realistic DeepFake setting, using the FakeAVCeleb dataset, highlighting the limitations of existing approaches when faced with semantic mismatch data. We further introduce three RARV-SMM variants that expose distinct architectural vulnerabilities as audio-visual divergence increases. We also propose a semantic reinforcement strategy that incorporates the semantic mismatch class and ImageBind embeddings to probe whether an explicit semantic coherence signal improves detection across architectures with different detection strategies, on FakeAVCeleb and LAVDF, contributing toward more realistic DeepFake detectors. The source code available at [https://github.com/](https://github.com/)[ ](https://github.com/)[sharayu-20/deepfake-semantic-mismatch](sharayu-20/deepfake-semantic-mismatch.)[.](sharayu-20/deepfake-semantic-mismatch.)  \n1. Introduction  \nAudio-visual content has become a critical medium for identity verification and political communication, raising concerns about the integrity of biometric systems that rely on facial/body appearance, voice, and lip dynamics [26, 27, 28, 22, 21, 11] . The rapid proliferation of DeepFake technologies substantially elevates associated risks, modern sys-  \nFigure 1 . Illustration of the proposed threat: a four-class DeepFake detector inspects authentic video from one event and authentic audio from another, finds no synthesis artifacts in either stream, and classifies the as safe, accepting potentially misleading content as genuine. This paper shows a vulnerability that semantic-aware detectors are better positioned to address, which is the basis of our proposal.  \ntems synthesize realistic faces [23, 30], clone voices [17], and generate lip-synchronized forgeries [24], enabling targeted disinformation, identity impersonation, and disbelief in audio-visual evidence [22, 11] . Detection methods have progressed from single modality artifact analysis [25, 16, 15] toward four-class audio-visual formulations [32, 18, 34] that discriminate between Real Audio– Real Video (RARV), Real Audio–Fake Video (RAFV), Fake Audio–Real Video (FARV), and Fake Audio–Fake Video (FAFV), providing greater forensic interpretability  \nthan binary systems. However, all existing approaches share a fundamental unaddressed assumption, that all real audio– real video samples are semantically consistent.  \nThis assumption fails in a class of realistic scenarios more dangerous than signal-level manipulations. We define semantic mismatch as the condition in which authentic audio and authentic video, each individually genuine, originate from different events, contexts, or narratives, such that their combined presentation conveys a fa","cbCaij1htmq6FRvr","https://ap.wps.com/l/cbCaij1htmq6FRvr","pdf",14290276,1,10,"English","en",105,"# Introduction\n## Semantic mismatch problem and threat model\n## Proposed evaluation setup and contributions","[{\"question\":\"What is semantic mismatch in the context of DeepFake detection?\",\"answer\":\"Semantic mismatch occurs when authentic audio and authentic video are each genuine but come from different events, contexts, or narratives, creating a misleading combined presentation without any synthesis artifacts.\"},{\"question\":\"How does the paper extend existing four-class audio-visual formulations?\",\"answer\":\"It introduces a new evaluation class, Real Audio–Real Video with Semantic Mismatch (RARV-SMM), explicitly modeling semantic-level inconsistency between authentic modalities.\"},{\"question\":\"What do the experiments show about state-of-the-art models under semantic mismatch?\",\"answer\":\"The study finds limitations in existing approaches when semantic mismatch is present, and it demonstrates how performance and robustness vary across models and experimental settings, especially as audio-visual divergence increases.\"}]",1784195981,25,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":85,"head_meta":87,"extra_data":89,"updated_unix":27},"are-deepfakes-realistic-enough-exploring-semantic-mismatch-as-a-novel-challenge","",{"@graph":35,"@context":84},[36,53,67],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/are-deepfakes-realistic-enough-exploring-semantic-mismatch-as-a-novel-challenge/84486/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":61,"encodingFormat":60,"isAccessibleForFree":62,"interactionStatistic":63},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-16",true,{"@type":64,"interactionType":65,"userInteractionCount":4},"InteractionCounter",{"@type":66},"ViewAction",{"@type":68,"mainEntity":69},"FAQPage",[70,76,80],{"name":71,"@type":72,"acceptedAnswer":73},"What is semantic mismatch in the context of DeepFake detection?","Question",{"text":74,"@type":75},"Semantic mismatch occurs when authentic audio and authentic video are each genuine but come from different events, contexts, or narratives, creating a misleading combined presentation without any synthesis artifacts.","Answer",{"name":77,"@type":72,"acceptedAnswer":78},"How does the paper extend existing four-class audio-visual formulations?",{"text":79,"@type":75},"It introduces a new evaluation class, Real Audio–Real Video with Semantic Mismatch (RARV-SMM), explicitly modeling semantic-level inconsistency between authentic modalities.",{"name":81,"@type":72,"acceptedAnswer":82},"What do the experiments show about state-of-the-art models under semantic mismatch?",{"text":83,"@type":75},"The study finds limitations in existing approaches when semantic mismatch is present, and it demonstrates how performance and robustness vary across models and experimental settings, especially as audio-visual divergence increases.","https://schema.org",{"og:url":51,"og:type":86,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":88,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":91},[92,96,100,104,109,114,119,122,127,130,133],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":93,"show_sort_weight":94,"slug":95},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":97,"show_sort_weight":98,"slug":99},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":101,"show_sort_weight":102,"slug":103},"Exam",70,"exam",{"id":105,"doc_module":4,"doc_module_name":45,"category_name":106,"show_sort_weight":107,"slug":108},5,"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":45,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":45,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":21,"doc_module":4,"doc_module_name":45,"category_name":131,"show_sort_weight":21,"slug":132},"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":45,"category_name":135,"show_sort_weight":105,"slug":136},19,"General","general"]