[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84102-en":3,"doc-seo-84102-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},84102,1099514067438,"River Wang","https://ap-avatar.wpscdn.com/avatar/100002539ee87300030?x-image-process=image/resize,m_fixed,w_180,h_180&k=1780474512215547542",8,"Research & Report","Generalized Synthetic Image Detection with Enhanced RGB-Noise Representation Learning","Large-scale generative models have intensified the spread of highly deceptive synthetic images, making generalized synthetic image detection a pressing need. Existing forensic methods often fail under cross-model shifts and real-world degradations because they depend on single-domain features and binary classification losses that produce fragile decision boundaries. We propose RNSIDNet, a dual-branch forensic framework using an attention-refined CLIP backbone, FiLM-based RGB-to-noise modulation, and Hard Sample-aware Contrastive Learning to improve margins and robustness across eight benchmarks.","Generalized Synthetic Image Detection with Enhanced RGB-Noise Representation Learning  \nZhen Lia , Gang Caoa,b,∗ , Tian Zhanga , Lifang Yuc and Shaowei Wengd  \na School of Computer and Cyber Sciences, Communication University of China, Beijing, 100024, China b School of Information Engineering, Changsha Medical University, Changsha, 410219, China  \nc Department of Information Engineering, Beijing Institute of Graphic Communication, Beijing, 100026, China  \nd Fujian Provincial Key Laboratory of Big Data Mining and Applications, Fujian University of Technology, Fuzhou, 350118, China  \nARTICLE INFO  \nKeywords:  \nImage forensics  \nAI-generated image Synthetic image detection Contrastive learning Feature fusion  \n7 Jul 2026  \nAB STRACT  \nThe rapid advancement of large-scale generative models has accelerated the spread of highly deceptive AI-generated images, making generalized synthetic image detection a critical imperative. Existing forensic networks often struggle with cross-model generalization and realworld degradations due to their reliance on single-domain representations and conventional binary classification optimization. To overcome these limitations, we propose RNSIDNet, a novel forensic framework that achieves robust detection through enhanced RGB-Noise representation learning. Specifically, our method employs a dual-branch architecture where global RGB semantics, extracted by an attention-refined CLIP backbone, dynamically modulate highfrequency noise artifacts captured by Bayar convolutions via a Feature-wise Linear Modulation (FiLM) module. To further enhance the learned representations, we design a Hard Sample-aware Contrastive Learning (HSCL) strategy. By explicitly penalizing challenging training samples, HSCL reshapes the latent feature space to maximize the discriminative margin between pristine and synthetic domains. Extensive experiments across eight public benchmark datasets verify that our model achieves state-of-the-art performance, delivering superior generalization ability, robustness, and computational efficiency. Code and dataset will be publicly available on [https:](https:)//[github.com/multimediaFor/RNSIDNet](github.com/multimediaFor/RNSIDNet).  \n1. Introduction  \nThe rapid evolution of deep learning techniques, especially large-scale generative models [13, 17, 37, 33], has  \n[ cs .CV]  \narXiv :2607 .06354v1  \nsignificantly advanced image synthesis while simultaneously facilitating the spread of deceptive AIGC content and deepfakes. In response, researchers have developed various forensic methods to identify these forgeries within a learning-based framework [44, 29, 2, 26] .  \nWhile numerous detection methods have been proposed, they frequently encounter a dual bottleneck in feature representation and model optimization when facing unseen generative models or complex real-world degradations, limiting their real-world applicability. A primary limitation of existing detectors is their reliance on single-domain representation. Spatial models are highly prone to overfitting dataset-specific semantics, whereas frequency-based approaches degrade sharply under common image corruptions like JPEG compression or Gaussian blur [44, 11] . To capture comprehensive forensic traces, recent studies [47, 41] have attempted to fuse diverse feature representations. Unfortunately, most existing methods rely on naive fusion strategies, such as direct concatenation or element-wise addition. These rigid operations fail to capture the complex contextual dependencies between different modalities. Consequently, such ineffective integration ultimately forces heterogeneous features to interfere with each other rather than achieve synergy, where dominant semantic information may obscure subtle frequency-domain artifacts.  \nBeyond representation, traditional optimization objectives further constrain detection capabilities. Most detectors [44, 29, 46] frame synthetic image detection as a standard binary classification task supervis","cbCaiieBkQkrk3aE","https://ap.wps.com/l/cbCaiieBkQkrk3aE","pdf",4563381,4,1,17,"English","en",105,"# Introduction\n## Problem: cross-model generalization and real-world degradations\n## Limitation: single-domain representation and naive fusion\n## Limitation: binary classification objectives\n## Proposed approach: RNSIDNet","[{\"question\":\"What problem does RNSIDNet address in synthetic image detection?\",\"answer\":\"RNSIDNet targets the difficulty of generalizing detection across unseen generative models and under real-world image degradations, which limits many existing forensic methods in practice.\"},{\"question\":\"How does RNSIDNet combine RGB semantics with noise artifacts?\",\"answer\":\"It uses a dual-branch architecture where global RGB semantics from an attention-refined CLIP backbone condition a noise branch. A FiLM module generates scale and bias to modulate high-frequency noise features extracted with Bayar convolutions.\"},{\"question\":\"What role does Hard Sample-aware Contrastive Learning (HSCL) play?\",\"answer\":\"HSCL explicitly emphasizes challenging training samples by penalizing easy or uninformative pairs less. This reshapes the feature space to enlarge the discriminative margin between pristine and synthetic domains.\"}]",1784192807,43,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"generalized-synthetic-image-detection-with-enhanced-rgb-noise-representation-learning","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/generalized-synthetic-image-detection-with-enhanced-rgb-noise-representation-learning/84102/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does RNSIDNet address in synthetic image detection?","Question",{"text":75,"@type":76},"RNSIDNet targets the difficulty of generalizing detection across unseen generative models and under real-world image degradations, which limits many existing forensic methods in practice.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does RNSIDNet combine RGB semantics with noise artifacts?",{"text":80,"@type":76},"It uses a dual-branch architecture where global RGB semantics from an attention-refined CLIP backbone condition a noise branch. A FiLM module generates scale and bias to modulate high-frequency noise features extracted with Bayar convolutions.",{"name":82,"@type":73,"acceptedAnswer":83},"What role does Hard Sample-aware Contrastive Learning (HSCL) play?",{"text":84,"@type":76},"HSCL explicitly emphasizes challenging training samples by penalizing easy or uninformative pairs less. This reshapes the feature space to enlarge the discriminative margin between pristine and synthetic domains.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]