[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84897-en":3,"doc-seo-84897-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},84897,1099514068035,"Ezra","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","Vendorbench-100: A Unified Cross-Paradigm Benchmark for Deepfake Image Detection","Deepfake image detection is currently delivered through three distinct paradigms—commercial detection APIs, zero-shot vision-language models, and open-source detectors—yet they are rarely assessed under a shared protocol, limiting meaningful comparison. VendorBench-100 evaluates 36 representative models using a single adversarial 100-image corpus, a unified output schema, and a common evaluation framework. It ranks models primarily by Matthews correlation coefficient under intentional class imbalance, with ROC-AUC reported for threshold-independent ranking ability.","VENDORBENCH-100: A UNIFIED CROSS-PARADIGM BENCHMARK FOR DEEPFAKE IMAGE DETECTION  \nSharayu N. Deshmukh  \nUniversidade da Beira Interior Covilha, Portugal [d.sharyu.nilesh@ubi.pt](d.sharyu.nilesh@ubi.pt)  \nMd Rashidunnabi  \nUniversidade da Beira Interior Covilha, Portugal [md.rashidunnabi@ubi.pt](md.rashidunnabi@ubi.pt)  \nNelton Tiago Gemo  \nUniversidade da Beira Interior Covilha, Portugal [nelton.gemo@ubi.pt](nelton.gemo@ubi.pt)  \narXiv :2607 .06254v 1 [ cs .CV] 7 Jul 2026  \nKurundkar G. D.  \nDepartment of Computer Science Shri Guru Buddhi Swami College Purna, India  \n[gajanan.kurundkar@gmail.com](gajanan.kurundkar@gmail.com)  \nMahamune M. R.  \nSchool of Computational Science Swami Ramanand Tirtha Marathwada University Nanded, India [mohnish.mahamune@gmail.com](mohnish.mahamune@gmail.com)  \nNilesh K. Deshmukh  \nSchool of Computational Science  \nSwami Ramanand Tirtha Marathwada University  \nNanded, India  \n[nileshkd.srt@gmail.com](nileshkd.srt@gmail.com)  \nJuly 8, 2026  \nABSTRACT  \nDeepfake image detection is currently served by three fundamentally different paradigms: commercial APIs, zero-shot vision-language models (LLMs), and open-source detectors. Despite their widespread use, these paradigms are rarely evaluated under a common protocol, making direct comparison difficult. We introduce VendorBench-100, a cross-paradigm benchmark that evaluates 36 representative models using a single adversarial 100-image corpus, a unified output schema, and a common evaluation framework. To ensure reliable assessment under the corpus’s intentional class imbalance, models are ranked primarily by the Matthews correlation coefficient (MCC), with ROC-AUC reported as a threshold-independent measure of ranking ability. Rather than maximizing dataset size, VendorBench-100 emphasizes challenging real-world scenarios through a curated taxonomy of eight edge-case families, including face swaps, text-to-video stills, AI photo edits, avatar compositing, opaque-provenance images, and compressed research frames. Our evaluation shows that commercial APIs achieve the strongest median performance, followed by vision LLMsand open-source detectors. However, individual open-source models remain competitive with the best vision LLMs. More importantly, we identify a consistent divergence between ranking ability (ROC-AUC) and operating-point quality (MCC), demonstrating that strong score discrimination does not necessarily produce reliable default-threshold decisions. This metric disagreement, rather than any single leaderboard ranking, is the central finding of the benchmark. We release the complete evaluation framework and benchmark results to support reproducible future research. The source code and data are available at: [https://github.com/sharayu-20/vendorbench-100](https://github.com/sharayu-20/vendorbench-100)  \n[Keywords:](Keywords: deepfake detection)[ deepfake detection](Keywords: deepfake detection); [AI-generated image detection](AI-generated image detection) ; [cross-paradigm benchmarking](cross-paradigm benchmarking) ; [vision](vision)  \nlanguage models; open-source detectors; Matthews correlation coefficient; reproducibility  \nA PREPRINT-JULY 8, 2026  \n1 Introduction  \nThe past several years have seen an extraordinary acceleration in the quality and accessibility of synthetic media. What once required specialized expertise and hand-tuned generative adversarial networks can now be accomplished by an ordinary user with a consumer-facing web application, a short text prompt, or a single uploaded photograph. Diffusionbased generators produce photorealistic faces and scenes frequently indistinguishable from genuine photographs under casual inspection, while face-swapping and avatar tools allow a person’s likeness to be transplanted into arbitrary video or image content in minutes. This democratization has a well-documented dark side: deepfakes and AI-generated imagery have been implicated in large-scale financial fraud carried out through impersonated","cbCaijG7WMIVx71g","https://ap.wps.com/l/cbCaijG7WMIVx71g","pdf",1569669,3,1,22,"English","en",105,"# Introduction\n## Problem: lack of cross-paradigm evaluation\n## Need for a unified benchmark\n## VendorBench-100 overview\n# Benchmark Design\n## Single adversarial 100-image corpus\n## Unified output schema and evaluation framework\n## Metrics: MCC and ROC-AUC\n# Edge-Case Taxonomy\n## Eight edge-case families\n# Results and Findings\n## Commercial APIs vs vision LLMs vs open-source detectors\n## Metric disagreement: ROC-AUC vs MCC\n# Reproducibility and Release","[{\"question\":\"Why is a unified benchmark needed for deepfake image detection?\",\"answer\":\"Commercial APIs, vision-language models used as zero-shot detectors, and open-source detectors are rarely evaluated under the same protocol. This makes direct comparison difficult for practitioners deciding how to allocate detection resources.\"},{\"question\":\"How does VendorBench-100 evaluate and rank models?\",\"answer\":\"Models are ranked primarily using the Matthews correlation coefficient (MCC) to handle the corpus’s intentional class imbalance. ROC-AUC is also reported as a threshold-independent measure of ranking ability.\"},{\"question\":\"What key finding does VendorBench-100 report about evaluation metrics?\",\"answer\":\"The benchmark identifies a consistent divergence between ranking ability (ROC-AUC) and operating-point quality (MCC). Strong score discrimination does not necessarily yield reliable decisions at a default threshold.\"}]",1784199182,55,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"vendorbench-100-a-unified-cross-paradigm-benchmark-for-deepfake-image-detection","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/vendorbench-100-a-unified-cross-paradigm-benchmark-for-deepfake-image-detection/84897/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-21","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why is a unified benchmark needed for deepfake image detection?","Question",{"text":75,"@type":76},"Commercial APIs, vision-language models used as zero-shot detectors, and open-source detectors are rarely evaluated under the same protocol. This makes direct comparison difficult for practitioners deciding how to allocate detection resources.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does VendorBench-100 evaluate and rank models?",{"text":80,"@type":76},"Models are ranked primarily using the Matthews correlation coefficient (MCC) to handle the corpus’s intentional class imbalance. ROC-AUC is also reported as a threshold-independent measure of ranking ability.",{"name":82,"@type":73,"acceptedAnswer":83},"What key finding does VendorBench-100 report about evaluation metrics?",{"text":84,"@type":76},"The benchmark identifies a consistent divergence between ranking ability (ROC-AUC) and operating-point quality (MCC). Strong score discrimination does not necessarily yield reliable decisions at a default threshold.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]