[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86032-en":3,"doc-seo-86032-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86032,1099514067438,"River Wang","https://ap-avatar.wpscdn.com/avatar/100002539ee87300030?x-image-process=image/resize,m_fixed,w_180,h_180&k=1780474512215547542",8,"Research & Report","Detecting AI-Generated Video: A Vision-Language Dual-View Survey","AI-generated videos (AIGC-V) are reaching cinematic realism, making traditional artifact-based detection insufficient. The survey reframes AIGC-V detection as Factual Fidelity Verification: whether depicted events, entities, and physical processes match real-world facts. It introduces a Vision-Language Dual-View taxonomy that organizes 221 reviewed works into a hierarchical four-layer framework, from intrinsic cue analysis to spatiotemporal consistency, cross-modal reasoning, and language-guided world-level verification, concluding with challenges and directions for robust, explainable, and trustworthy detection.","arXiv :2607 . 10787v1 [ cs .CV] 12 Jul 2026  \nDetecting AI-Generated Video: A Vision-Language Dual-View Survey  \nDylan Xinming Hou1, Juntian Zhang2, Xu Gu2,  \nYichen Wu3, Nils Lukas1, Gus Xia1, Xiuying Chen1, Yuhan Liu1†  \n1MBZUAI, 2 Gaoling School of Artificial Intelligence, Renmin University of China, 3 Harvard University  \n[dyxhou@gmail.com](dyxhou@gmail.com) , {nils.lukas, gus .xia, xiuying.chen, [yuhan.liu}@mbzuai.ac.ae](yuhan.liu}@mbzuai.ac.ae) ,{zhangjuntian, [guxu}@ruc.edu.cn](guxu}@ruc.edu.cn) , [yiwu6@mgh.harvard.edu](yiwu6@mgh.harvard.edu)  \nThe evolving realism of AI-generated Videos (AIGC-V) is rapidly rendering traditional artifact-centric detection insufficient, necessitating a paradigm shift from low-level inspection to high-level semantic verification. This paper presents a comprehensive survey of AIGC-V detection, reframing the task as Factual Fidelity Verification, which asks whether the events, entities, and physical processes depicted in a video are consistent with real-world facts. To systematize this rapidly evolving field, we propose a Vision-Language Dual-View taxonomy that organizes existing methods into a hierarchical, four-layer landscape, spanning intrinsic cue analysis, spatiotemporal consistency modeling, cross-modal consistency reasoning, and language-guided world-level reasoning. This dual-view framing highlights a fundamental transition from artifact matching in traditional deepfake detection to evidence-based semantic verification enabled by vision-language models and agentic reasoning pipelines. Based on a systematic review of 221 works, we synthesize AIGC-V generation paradigms, survey the landscape of detection methods, and review evaluation metrics and benchmarks in line with proposed views. Finally, we discuss current challenges and identify promising directions toward robust, explainable, and trustworthy detection.  \nDate: July 14, 2026  \n§ Github: [https://github.com/dxhou/AI-Generated-Video-Detection](https://github.com/dxhou/AI-Generated-Video-Detection)[ ](https://github.com/dxhou/AI-Generated-Video-Detection)Ñ Homepage: [https://aigcvdetection.github.io/](https://aigcvdetection.github.io/)  \n1 Introduction  \nThe rapid evolution of video generation models, exemplified by Sora 2 (OpenAI, 2024b), Veo 3 (Google DeepMind, 2025), and Seedance 2.0 (ByteDance Seed, 2026), is fundamentally reshaping the information landscape. Unlike early deepfakes dominated by localized face swapping (Wang et al., 2024b), modern AI-Generated Content-Videos (AIGC-V) have achieved cinematic fidelity with coherent narratives, increasingly blurring the boundary between synthesized fiction and captured reality. This technological leap destabilizes the foundational trust in video evidence, traditionally relied upon to verify \"who, when, where, and what\" (Ho et al., 2022a) .  \n†Corresponding author: Yuhan Liu([yuhan.liu@mbzuai.ac.ae](yuhan.liu@mbzuai.ac.ae))  \nFigure 1 An example of the AIGC-V detection pipeline under our approach, illustrating AIgenerated videos from traditional methods or text-to-video prompts, detection from the visual and language views, and outputs at different levels.  \nAs generation paradigms shift from local manipulation to end-to-end synthesis, such as text-to-video, traditional detection methods (Wang et al., 2025d; Ma et al., 2025) face a critical bottleneck. Early detection systems primarily relied on low-level visual artifacts, such as blending boundaries. However, advanced diffusion models and transformers (Ho et al., 2022b; OpenAI, 2024a) can now produce visually high-fidelity videos. This paradigm shift necessitates a new detection landscape, shifting from perceptual inspection (checking for visual artifacts) to cognitive reasoning (checking for semantic and factual violations) . The surge of Vision-Language Models (VLMs)(Zhang et al., 2025d,c; Wang et al., 2026) and agentic frameworks(Liu et al., 2024b, 2025g,f; Hou et al., 2024) offers a promising pathway, enabling detectors to ","cbCailvON2SlT30B","https://ap.wps.com/l/cbCailvON2SlT30B","pdf",3792395,3,1,51,"English","en",105,"# Introduction\n## Vision-Language Dual-View Taxonomy\n## Factual Fidelity Verification Framework\n## Survey Coverage and Structure\n# Challenges and Future Directions","[{\"question\":\"What problem does the survey address in detecting AI-generated videos?\",\"answer\":\"It addresses the inadequacy of traditional artifact-centric detection as AI-generated videos become visually realistic and harder to distinguish using low-level cues alone.\"},{\"question\":\"How does the survey redefine video detection?\",\"answer\":\"It reframes detection as Factual Fidelity Verification, checking whether the video’s events, entities, and physical processes are consistent with real-world facts and physical laws.\"},{\"question\":\"What is the proposed Vision-Language Dual-View taxonomy?\",\"answer\":\"It organizes existing methods into four layers: intrinsic cue analysis, spatiotemporal consistency, cross-modal consistency reasoning, and language-guided world-level reasoning to support evidence-based semantic verification.\"}]",1784207961,129,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"detecting-ai-generated-video-a-vision-language-dual-view-survey","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/detecting-ai-generated-video-a-vision-language-dual-view-survey/86032/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the survey address in detecting AI-generated videos?","Question",{"text":75,"@type":76},"It addresses the inadequacy of traditional artifact-centric detection as AI-generated videos become visually realistic and harder to distinguish using low-level cues alone.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the survey redefine video detection?",{"text":80,"@type":76},"It reframes detection as Factual Fidelity Verification, checking whether the video’s events, entities, and physical processes are consistent with real-world facts and physical laws.",{"name":82,"@type":73,"acceptedAnswer":83},"What is the proposed Vision-Language Dual-View taxonomy?",{"text":84,"@type":76},"It organizes existing methods into four layers: intrinsic cue analysis, spatiotemporal consistency, cross-modal consistency reasoning, and language-guided world-level reasoning to support evidence-based semantic verification.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]