[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82247-en":3,"doc-seo-82247-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},82247,962075114765,"Quinn","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","YeTI: You Only Need Two Noisy Images for Real-World sRGB Noise Generation","Real-world sRGB image denoising is hindered by nonlinear sensor-noise behavior and the difficulty of obtaining aligned clean–noisy pairs. Supervised denoisers can overfit to limited paired datasets, and self-supervised approaches still require sufficiently diverse noisy observations. This work introduces YeTI, a clean-image-free, metadata-free framework for realistic noise synthesis from only two noisy observations of the same scene. A Reconstruction Autoencoder disentangles structure and noise, while a one-step Conditional Diffusion Transformer models latent noise under consistency objectives. Experiments validate effectiveness on SIDD, generalization on SIDD+, MAI2021, and SID, and downstream denoising gains on DND.","YeTI: You Only Need Two Noisy Images for Real-World sRGB Noise Generation  \narXiv :2607 .09193v1 [ cs .CV] 10 Jul 2026  \nJaekyun Ko 1 ,2∗, Byung Wan Lim 1∗, Soomin Lee 1 , Dongjin Kim 1 , and  \nTae Hyun Kim 1†  \n1 Department of Computer Science, Hanyang University {pook0612,min001017,dongjinkim,[taehyunkim}@hanyang.ac.kr](taehyunkim}@hanyang.ac.kr)  \n2 Mobile Experience (MX) Division, Samsung Electronics  \n[jkko1124.ko@samsung.com](jkko1124.ko@samsung.com)  \n\n|  | \u003Cbr> |\n| --- | --- |\n\nFig. 1: Noise Generation Comparison. (a) Conventional clean-image-dependent approaches. (b) Our clean-image-free approach.  \nAbstract. Real-world sRGB image denoising remains challenging due to the nonlinear characteristics of sensor noise and the difficulty of acquiring aligned clean-noisy image pairs. Supervised denoisers often overfit to limited paired datasets, while self-supervised methods still depend on sufficiently diverse noisy observations. These limitations motivate scalable noise synthesis methods that can model real-world noise without clean ground truth or camera metadata. We propose YeTI, a real-world sRGB noise generation framework that learns from only two noisy observations of the same scene. YeTI uses a Reconstruction Autoencoder to disentangle scene structure and noise characteristics, and models the latent noise distribution with a one-step Conditional Diffusion Transformer trained using consistency objectives. Given a single noisy input at inference time, YeTI generates realistic, signal-dependent noise while preserving the underlying scene content. Extensive experiments demonstrate the effectiveness of YeTI across real-world benchmarks. We evaluate noise generation on SIDD and further assess generalization on SIDD+, MAI2021, and SID, covering smartphone and diverse consumer-camera sensors. Downstream denoising results on DND further show that denoisers trained with YeTI-synthesized images achieve strong real-world performance, highlighting the practical value of clean-image-free and metadata-free noise generation. Code is available at: [https://github.com/ByungWanLim/YeTI-You-Only-Need](https://github.com/ByungWanLim/YeTI-You-Only-Need)Two-Noisy-Images-for-Real-World-sRGB-Noise-Generation  \n* Equal contribution. † Corresponding author.  \n2 J. Ko et al.  \nTable 1: Comparison between real-world sRGB noise modeling methods.  \n\n| Generative Model | No Need for Clean? Realistic Noise? No Need for Metadata? |  |  |\n| --- | --- | --- | --- |\n| (a) NeCA, NAFlow | ✗ | ✓ | ✗ |\n| (b) C2N | ✗ | ✗ | ✓ |\n| (c) SeNM-VAE | ✗ | ✓ | ✓ |\n| (d) YeTI (Ours) | ✓ | ✓ | ✓ |\n\n1 Introduction  \nImage denoising is a fundamental problem in computer vision and plays a critical role in various real-world applications such as photography [7,19], surveillance [63], and autonomous driving [27, 45, 52] . Despite the remarkable progress of deep learning-based denoisers [8,9,18,38,60,66,67,71], achieving robust performance in various imaging conditions remains a challenging problem.  \nIn particular, real-world sRGB denoising poses unique difficulties due to the complex characteristics of sensor noise and the nonlinear transformations applied during the image signal processing (ISP) pipeline [14, 15, 50] . The raw sensor noise becomes highly distorted after passing through demosaicing, tone mapping, gamma correction, and other camera-specific operations. Consequently, the noise distribution in the real-world sRGB domain deviates significantly from simple Poisson-Gaussian (PG) assumptions [12,34], making it difficult for supervised denoising models to generalize well to real-world settings.  \nMoreover, the scarcity of real-world training datasets leads to overfitting problems in supervised denoising methods. Although datasets such as SIDD [3] and MIDD [11] attempt to collect real-world noisy–clean image pairs, the process is labor-intensive and expensive. This process requires capturing more than a thousand short- and long-exposure images for each scene under f","cbCailYQgr42DsGB","https://ap.wps.com/l/cbCailYQgr42DsGB","pdf",3981078,1,30,"English","en",105,"# Introduction\n## Problem Background: Real-World sRGB Denoising Challenges\n## Proposed Solution: YeTI Framework\n### Reconstruction Autoencoder and Noise Disentanglement\n### One-step Conditional Diffusion Transformer for Latent Noise Modeling\n# Noise Generation Comparisons and Experimental Evaluation","[{\"question\":\"Why is real-world sRGB denoising especially challenging?\",\"answer\":\"Real-world sensor noise undergoes nonlinear transformations in the ISP pipeline, so the noise distribution in the sRGB domain deviates from simple Poisson-Gaussian assumptions. This makes robust generalization difficult across imaging conditions.\"},{\"question\":\"What does YeTI remove compared with prior sRGB noise modeling methods?\",\"answer\":\"YeTI eliminates the need for paired clean–noisy image pairs and the need for camera metadata, synthesizing realistic sRGB noise without requiring those inputs.\"},{\"question\":\"How does YeTI generate noise at training and inference time?\",\"answer\":\"YeTI learns from only two noisy burst observations during training and uses a single noisy image at inference to generate realistic, signal-dependent noise while preserving underlying scene content.\"}]",1784179134,76,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"yeti-you-only-need-two-noisy-images-for-real-world-srgb-noise-generation","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/yeti-you-only-need-two-noisy-images-for-real-world-srgb-noise-generation/82247/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why is real-world sRGB denoising especially challenging?","Question",{"text":75,"@type":76},"Real-world sensor noise undergoes nonlinear transformations in the ISP pipeline, so the noise distribution in the sRGB domain deviates from simple Poisson-Gaussian assumptions. This makes robust generalization difficult across imaging conditions.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What does YeTI remove compared with prior sRGB noise modeling methods?",{"text":80,"@type":76},"YeTI eliminates the need for paired clean–noisy image pairs and the need for camera metadata, synthesizing realistic sRGB noise without requiring those inputs.",{"name":82,"@type":73,"acceptedAnswer":83},"How does YeTI generate noise at training and inference time?",{"text":84,"@type":76},"YeTI learns from only two noisy burst observations during training and uses a single noisy image at inference to generate realistic, signal-dependent noise while preserving underlying scene content.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":21,"slug":121},"research-report",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":45,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":45,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":45,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]