[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86007-en":3,"doc-seo-86007-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":11,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},86007,1099514067415,"Rowan","https://ap-avatar.wpscdn.com/avatar/100002539d78ffe74a7?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779092875211072502",8,"Research & Report","Effective Synthetic Image Detection via Noise Residual Clustering","Rapid progress in generative AI has produced synthetic images that closely mimic real photographs, enabling misinformation and fraud. The work addresses blind, passive detection where supervised training on large labeled sets is costly and degrades on unseen generators. A training-free framework is proposed: Noiseprint++ extracts noise residual fingerprints, a frozen ViT derives multi-scale residual features, and adaptive weighted fusion prepares representations. Unsupervised K-Means is initialized from a few real samples’ clustering centers. Evaluations on four benchmarks show an average accuracy of 82.2% with strong generalization, especially for diffusion-based images, supported by ablation studies.","This is a preprint as arXiv: 2607 [cs.CV] July 2026.  \nEffective Synthetic Image Detection via Noise Residual Clustering  \nCaihui Yan 1, Gang Cao 1,*, Huawei Tian2, Zhen Li1, Yuhang Zhai1  \n1 School of Computer and Cyber Sciences, Communication University of China, Beijing 100024, China  \n2 People's Public Security University of China, Beijing 100038, China  \nAbstract—The rapid advancement of generative artificial intelligence (AI) has made synthetic images remarkably realistic, posing security threats such as misinformation and fraud. It is significant to detect the synthetic image in the manner of passive and blind image authentication. Most existing detectors rely on supervised training with large labeled datasets, leading to high costs and degraded performance on unknown generative models. To attenuate such deficiencies, we propose a training-free detection method. Specifically, noise residual fingerprints are first extracted by a simple yet effective pre-trained Noiseprint++ model. Then multi-scale features are further extracted from such residual by a frozen Vision Transformer (ViT), followed by adaptive weighted fusion. Only a few real image samples are used needed to initialize the clustering centers for unsupervised K-Means, distinguishing real and synthetic images without training. Extensive evaluations on four benchmark datasets show that our proposed scheme achieves an average accuracy of 82.2%, outperforming the state-of-the-art detectors on generalization ability. Superior performance is gained on the popular diffusion type of synthetic images, and the effectiveness of each module is validated by ablation studies. Source code will be publicly available at [https://github.com/multimediaFor/NoiseCluSID](https://github.com/multimediaFor/NoiseCluSID).  \nIndex Terms—Synthetic image detection, Noise residual, Feature fusion, Clustering, Training-free  \nI. INTRODUCTION  \nThe rapid advancement of generative artificial intelligence (AI) techniques has made AI-generated images and videos visually indistinguishable from real photographs, raising serious concerns on potential misuse of such synthetic images. Generative Adversarial Networks (GANs), such as ProGAN [1], StyleGAN [2], and StarGAN [3], can synthesize highly realistic facial images. More recently, diffusion models, including Stable Diffusion [4], DALL·E 2 [5] and Midjourney [6], have achieved superior generation quality and become mainstream. The resulting AI-generated images may be maliciously exploited to spread disinformation, manipulate public opinion, and undermine the security of digital  \n* Corresponding author: Gang Cao([gangcao@cuc.edu.cn](gangcao@cuc.edu.cn))  \ncontent. Therefore, it is significant to develop generalizable and robust blind detectors capable of distinguishing AI-generated images/videos from real ones [7-25] .  \nExisting synthetic image detection methods can be broadly categorized into supervised and trainingfree types. Supervised approaches train classifiers on large labeled datasets to distinguish real from synthetic images. Early efforts established strong baselines by fine-tuning off-the-shelf networks (Marra et al. [7]), and were later improved by architectural modifications such as inserting residual layers and removing early downsampling (Gragnaniello et al. [8]) . Durall et al. [9] and Frank et al. [10] showed that CNN-generated images contain detectable frequency artifacts, while Corvi et al. [11] extended detection to latent diffusion models. A notable milestone is Wang et al. [12], who designed a ResNet50-based detector with aggressive data augmentation that significantly boosted generalization and became a widely adopted benchmark. Recent works diverge into several directions. Leveraging vision-language models, UnivFD [13] trains a linear classifier on frozen CLIP [14] features, achieving promising cross-model generalization. CLIP-Bar [15] and C2P-CLIP [16] further enhance the exploitation of CLIP representations for this task.","cbCaijKTu068n8Rw","https://ap.wps.com/l/cbCaijKTu068n8Rw","pdf",3132190,1,17,"English","en",105,"# Introduction\n## Background on generative AI and detection needs\n## Related work: supervised detectors\n## Related work: training-free detectors","[{\"question\":\"Why is synthetic image detection difficult for unknown generative models?\",\"answer\":\"Supervised detectors rely on training with large labeled datasets, which leads to high cost and degraded performance when facing unseen generators. Their learned forensic signals may not generalize well across model types.\"},{\"question\":\"How does the proposed training-free method identify real versus synthetic images?\",\"answer\":\"It first extracts noise residual fingerprints using the pretrained Noiseprint++ model, then feeds residuals into a frozen Vision Transformer to obtain multi-scale features. After adaptive weighted fusion, unsupervised K-Means separates real and synthetic images using clustering centers initialized from only a few real samples.\"},{\"question\":\"What evidence supports the effectiveness of the method?\",\"answer\":\"Experiments on four benchmark datasets report an average accuracy of 82.2%, outperforming state-of-the-art detectors in generalization. Performance is especially strong on diffusion-type synthetic images, and ablation studies validate each module’s contribution.\"}]",1784207741,43,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"effective-synthetic-image-detection-via-noise-residual-clustering","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/effective-synthetic-image-detection-via-noise-residual-clustering/86007/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":11},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why is synthetic image detection difficult for unknown generative models?","Question",{"text":75,"@type":76},"Supervised detectors rely on training with large labeled datasets, which leads to high cost and degraded performance when facing unseen generators. Their learned forensic signals may not generalize well across model types.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the proposed training-free method identify real versus synthetic images?",{"text":80,"@type":76},"It first extracts noise residual fingerprints using the pretrained Noiseprint++ model, then feeds residuals into a frozen Vision Transformer to obtain multi-scale features. After adaptive weighted fusion, unsupervised K-Means separates real and synthetic images using clustering centers initialized from only a few real samples.",{"name":82,"@type":73,"acceptedAnswer":83},"What evidence supports the effectiveness of the method?",{"text":84,"@type":76},"Experiments on four benchmark datasets report an average accuracy of 82.2%, outperforming state-of-the-art detectors in generalization. Performance is especially strong on diffusion-type synthetic images, and ablation studies validate each module’s contribution.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]