[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83425-en":3,"doc-seo-83425-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83425,7971461741311,"Ophelia","https://ap-avatar.wpscdn.com/avatar/74000253aff267980c6?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779345379180704826",8,"Research & Report","HumanForge: 人本中心深伪视频基准及多智能体伪造理由","Rapid progress in video diffusion and temporal editing enables highly realistic human-centric forgeries, creating urgent challenges for digital forensics. Existing benchmarks largely target face-swapping or global text-to-video synthesis and miss key factors like human-object and human-human interactions and multimodal alignment. HumanForge introduces a unified, large-scale, multi-paradigm human-centric video forgery dataset with 18K+ high-fidelity segments. Gen2Anno, a LangGraph-based multiagent pipeline, produces structured omni-annotations and supports fine-grained benchmarks for zero-shot generalization and reasoning.","HumanForge: A Human-Centric Deepfake Video Benchmark with  \nMulti-Agent Forgery Rationales  \nWenbo Xu 1, * , Zhimin Chen 1, * , Xiaojie Liang 1, * , Hengrui Liu 1, * , Wei Lu 1†  \n1 School of Computer Science and Engineering, Sun Yat-sen University, Guangzhou 510006, China  \narXiv :2607 .08705v 1 [ cs .CV] 9 Jul 2026  \nAbstract  \nRapid advancements in video diffusion models and temporal editing tools have enabled the generation of highly realistic human-centric videos, posing unprecedented challenges to digital content forensics. Existing benchmarks primarily focus on either face-swapping or global text-to-video synthesis, overlooking the crucial dimensions of human-object/humanhuman interactions (HOI/HHI) and multi-modal alignment. To address these limitations, we introduce HumanForge, a unified, large-scale, and multi-paradigm human-centric video forgery dataset. To construct and annotate this dataset without labor-intensive manual labeling or hallucinated monolithic prompts, we propose Gen2Anno, a modular active multiagent pipeline built on LangGraph. Gen2Anno coordinates six specialized agents—ranging from source profiling to MoE-based reference analysis and closed-loop forensic verification—to generate over 18K high-fidelity video segments and produce structured, contrastive “omni-annotations” containing binary decisions, fine-grained artifact categories, and spatio-temporal localization. Extensive benchmarks using state-of-the-art traditional detectors and Large Multimodal Models (LMMs) demonstrate the significant challenges of zero-shot generalization and fine-grained reasoning on HumanForge. Code and dataset will be publicly released.  \nIntroduction  \nThe rapid advancement of generative artificial intelligence, particularly diffusion-based video generation and editing architectures, has made it possible to synthesize highly realistic human-centric videos. Technologies such as Wan2.1, CogVideoX, and LTX-Video can now generate intricate human motions, precise lip movements, and complex social interactions with impressive visual fidelity. However, this progress also lowers the barrier to creating sophisticated deepfakes, posing significant risks to information security, public trust, and social stability. Consequently, developing robust deepfake detection methods—especially those capable of identifying human-centric video forgeries—has become an urgent priority for the artificial intelligence community.  \n*These authors contributed equally.  \n†Corresponding author.  \nCopyright © 2026, Association for the Advancement of Artificial Intelligence ([www.aaai.org](www.aaai.org)). All rights reserved.  \nDespite the critical need, existing modern deepfake detection datasets and benchmarks suffer from several key limitations. First, while image-based explainable frameworks such as Veritas (Tan et al. 2026b) and AnomReason (Tanet al. 2026a) have made strides in identifying synthetic artifacts, they are restricted to static visual representations. Consequently, they fail to capture the complex temporal, kinematic, and audio-visual dynamics inherent to video forgeries.  \nSecond, contemporary video-centric benchmarks often suffer from narrow scenario coverage. For instance, AvatarShield (Xu et al. 2025) focuses primarily on talking avatars and audio-driven speech, while ActivityForensics (Bao et al. 2026) targets temporal activity localization. These benchmarks frequently overlook complex interactive scenarios, such as human-human or human-object interactions, focusing instead on isolated subjects. In real-world forensic scenarios, human activities are rarely isolated; understanding physical and spatial consistency during interactions is crucial for identifying sophisticated forgeries.  \nFinally, a major bottleneck lies in the annotation of these deepfake datasets. Existing automated annotation paradigms typically inspect the generated media in a vacuum, without access to the underlying generation metadata. This lack of context makes ","cbCaiak08roVJ4aC","https://ap.wps.com/l/cbCaiak08roVJ4aC","pdf",14038039,3,1,6,"English","en",105,"# Abstract\n# Introduction\n## Key Limitations of Existing Benchmarks\n## HumanForge Dataset and Scenarios\n## Gen2Anno Multi-Agent Annotation Pipeline","[{\"question\":\"HumanForge解决了现有深伪视频基准的哪些主要不足？\",\"answer\":\"它补足了以往基准在交互场景（人-物/人-人）、多模态对齐以及仅关注单一合成范式方面的不足，并提供更具语境的精细取证标注。\"},{\"question\":\"HumanForge的数据规模和覆盖的核心人本中心合成场景是什么？\",\"answer\":\"数据集包含18,000+合成视频，并系统覆盖四类场景：音频驱动唇同步、姿态驱动动作迁移、人际/人-物交互建模，以及语义驱动的文本到视频编辑。\"},{\"question\":\"Gen2Anno如何避免盲目标注带来的误差并生成“omni-annotations”？\",\"answer\":\"Gen2Anno将生成溯因信息与视觉分析结合，基于参考与检查代理构建“Expected State”，再进行闭环取证验证，输出二元决策、细粒度伪造类别与时空定位等结构化标注。\"}]",1784187517,15,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"humanforge-a-human-centric-deepfake-video-benchmark-with-multi-agent-forgery-rationales","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/humanforge-a-human-centric-deepfake-video-benchmark-with-multi-agent-forgery-rationales/83425/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"HumanForge解决了现有深伪视频基准的哪些主要不足？","Question",{"text":75,"@type":76},"它补足了以往基准在交互场景（人-物/人-人）、多模态对齐以及仅关注单一合成范式方面的不足，并提供更具语境的精细取证标注。","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"HumanForge的数据规模和覆盖的核心人本中心合成场景是什么？",{"text":80,"@type":76},"数据集包含18,000+合成视频，并系统覆盖四类场景：音频驱动唇同步、姿态驱动动作迁移、人际/人-物交互建模，以及语义驱动的文本到视频编辑。",{"name":82,"@type":73,"acceptedAnswer":83},"Gen2Anno如何避免盲目标注带来的误差并生成“omni-annotations”？",{"text":84,"@type":76},"Gen2Anno将生成溯因信息与视觉分析结合，基于参考与检查代理构建“Expected State”，再进行闭环取证验证，输出二元决策、细粒度伪造类别与时空定位等结构化标注。","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,114,119,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":111,"show_sort_weight":112,"slug":113},"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]