[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84396-en":3,"doc-seo-84396-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},84396,7971461741311,"Ophelia","https://ap-avatar.wpscdn.com/avatar/74000253aff267980c6?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779345379180704826",8,"Research & Report","Do Transformations Reveal the Truth? Generative Residual Learning for Generalized AI-Generated Image Detection","The rapid rise of generative AI enables highly realistic deepfake and AI-generated images, driving risks such as misinformation, digital identity theft, fraud, and manipulation of public opinion. Reliable AI-generated image detection remains difficult because diverse generation pipelines leave subtle, variable artifacts and often fail under unseen generators. This work introduces GenRes, a generative residual learning framework using relational modeling between original and transformed samples via a neural tensor network. For multiple transformations, GenRes++ adds learnable attention aggregation, improving cross-domain generalization on benchmark datasets.","Do Transformations Reveal the Truth? Generative Residual Learning for Generalized AI-Generated Image Detection  \nKutub Uddin University of Michigan Flint, Michigan, USA  \n[kutub@umich.edu](kutub@umich.edu)  \nNusrat Tasnim Korea Aerospace University Goyang, South Korea  \n[tasnim.nishu70@kau.kr](tasnim.nishu70@kau.kr)  \nAwais Khan University of Michigan Flint, Michigan, USA  \n[mawais@umich.edu](mawais@umich.edu)  \nMohammad Umar Farooq University of Michigan Flint, Michigan, USA  \n[mufarooq@umich.edu](mufarooq@umich.edu)  \nKhalid Malik University of Michigan Flint, Michigan, USA  \n[drmalik@umich.edu](drmalik@umich.edu)  \narXiv :2607 .08674v 1 [ cs .CV] 9 Jul 2026  \nAbstract  \nThe rapid advancement of generative AI has enabled the creation of highly realistic deepfake media, posing significant threats, including misinformation, digital identity theft, fraud, and manipulation of public opinion. AI-generated image (AIGI) detection is reliably challenging due to the diversity of generative methods and the subtle artifacts they leave behind. In this work, we propose GenRes, a novel framework for generative residual learning via a neural tensor network, which models fine-grained relational features between original and transformed samples to enhance generalization. To address scenarios involving multiple generative transformations, we introduce GenRes++, which employs a learnable attention mechanism to aggregate relational features across multiple transformed samples and enables the model to focus on the most informative cues. Both models leverage PE-Core as a feature extractor, providing generalized and semantically rich embeddings that improve cross-domain performance and enable the detection of AIGI generated by unseen methods. Comprehensive experiments on multiple benchmark datasets demonstrate that the proposed GenRes++ approach outperforms existing methods.  \n1. Introduction  \nAdvances in generative adversarial networks (GANs) [1], diffusion models [2], and hybrid [3] synthesis pipelines have made AIGIs increasingly indistinguishable from real photographs, fueling misinformation, identity fraud, and deepfake-related financial losses now reaching hundreds of millions of dollars annually [4, 5, 6, 7] . Detecting such content reliably is therefore an urgent and practically consequential problem. Existing detectors fall into two paradigms: artifact-driven methods exploit  \nFigure 1 . Comparison between traditional AIGI detection and the proposed GenRes++ framework. (a) Traditional approaches encode a single image and perform direct detection, relying on isolated features, thereby limiting their ability to capture subtle generative artifacts and weakening cross-generator generalization. (b) In contrast, our approach generates multiple transformed variants and learns relational residuals between the original and transformed images. By exploiting how generative artifacts interact under transformations, the model captures more discriminative properties, leading to improved generalization.  \ngenerator-specific fingerprints [8, 9, 10, 11, 12, 13], but degrade sharply on unseen generators, while representationlearning methods project images into forgery-sensitive CLIP spaces [14, 15, 16, 17] yet still encode each image independently, neglecting the relational structure induced by generative processing. As a result, neither paradigm achieves generalization across the rapidly expanding diversity of generative models.  \nReal and synthetic images respond differently when processed by a second generative model (restoration, superresolution, enhancement, or denoising) [18, 19] . Natural images produce outputs broadly consistent with their rich, high-frequency content, whereas AIGIs exhibit characteristic residual discrepancies as the transform interacts with the original generator’s statistical biases [20, 21, 22, 23, 24] . This occurs because an AIGI already embodies the distributional priors of its source generator. When a second generative mod","cbCaivnFNzv1uouZ","https://ap.wps.com/l/cbCaivnFNzv1uouZ","pdf",3890942,5,1,12,"English","en",105,"# Introduction\n## Background and challenges in AIGI detection\n## Proposed framework: GenRes and GenRes++\n## Model design and contributions","[{\"question\":\"Why is generalized AI-generated image detection difficult with existing methods?\",\"answer\":\"Existing artifact-driven detectors rely on generator-specific fingerprints and degrade on unseen generators, while representation-learning approaches encode images independently and miss relational structure created by generative processing.\"},{\"question\":\"What idea motivates GenRes and GenRes++?\",\"answer\":\"Natural and synthetic images respond differently to a second generative transformation; synthetic images exhibit characteristic discrepancies termed generative residuals, so the method models relationships between an original image and transformed variants rather than analyzing images in isolation.\"},{\"question\":\"How do GenRes and GenRes++ use transformations to improve generalization?\",\"answer\":\"GenRes models the relational structure between original and transformed samples using a neural tensor network to capture bilinear interactions. GenRes++ extends this with learnable attention aggregation across multiple transformed samples, focusing on the most informative residual cues.\"}]",1784195301,30,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"do-transformations-reveal-the-truth-generative-residual-learning-for-generalized-ai-generated-image-detection","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/do-transformations-reveal-the-truth-generative-residual-learning-for-generalized-ai-generated-image-detection/84396/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"Why is generalized AI-generated image detection difficult with existing methods?","Question",{"text":76,"@type":77},"Existing artifact-driven detectors rely on generator-specific fingerprints and degrade on unseen generators, while representation-learning approaches encode images independently and miss relational structure created by generative processing.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"What idea motivates GenRes and GenRes++?",{"text":81,"@type":77},"Natural and synthetic images respond differently to a second generative transformation; synthetic images exhibit characteristic discrepancies termed generative residuals, so the method models relationships between an original image and transformed variants rather than analyzing images in isolation.",{"name":83,"@type":74,"acceptedAnswer":84},"How do GenRes and GenRes++ use transformations to improve generalization?",{"text":85,"@type":77},"GenRes models the relational structure between original and transformed samples using a neural tensor network to capture bilinear interactions. GenRes++ extends this with learnable attention aggregation across multiple transformed samples, focusing on the most informative residual cues.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,110,115,120,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":29,"slug":121},"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":20,"slug":137},19,"General","general"]