[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86290-en":3,"doc-seo-86290-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86290,1099514068365,"Aurelia","https://ap-avatar.wpscdn.com/avatar/10000253d8d9f28188e?_k=1776742907772140068",8,"Research & Report","Feature-Space Guided Diffusion for Realistic Ultrasound Image Synthesis","Conditional diffusion models can generate anatomically plausible medical ultrasound images, yet anatomical alignment alone does not ensure realistic B-mode appearance. Standard pipelines either condition generative architectures on anatomical masks or employ guidance mechanisms reinforcing the same anatomical signal, while ultrasound realism depends on acquisition-linked properties such as speckle texture, tissue contrast, and attenuation. Using a frozen ultrasound foundation model, Feature-Space Candidate Guidance (FSCG) reduces a representation-space gap via training-free k-NN feature correction and candidate selection.","arXiv :2607 . 11655v1 [ cs .CV] 13 Jul 2026  \nFeature-Space Guided Diffusion for Realistic Ultrasound Image Synthesis  \nMarina Domínguez 1 (􀀀) , Nélida Mirabet-Herranz 1 , and Valery Naranjo 1 ,2  \n1 Instituto Universitario de Investigación en Tecnología Centrada en el Ser Humano,  \nHUMAN-tech, Universitat Politècnica de València, Valencia, Spain [mdommar1@doctor.upv.es](mdommar1@doctor.upv.es)  \n2 Artikode Intelligence S.L. , Valencia, Spain  \nAbstract. Conditional diffusion models can generate anatomically plausible medical ultrasound (US) images, but anatomical plausibility alone does not ensure realistic B-mode appearance. Most US pipelines adapt standard generative architectures and condition them on anatomical masks, or use guidance mechanisms that reinforce the same anatomical signal. However, B-mode US images are shaped by acquisition-dependent properties such as speckle texture, tissue contrast, and attenuation. Using a frozen US foundation model, we show that standard conditional diffusion baselines remain separated from real images in representation space.  \nIn this work, we propose Feature-Space Candidate Guidance (FSCG), a training-free sampling strategy to reduce this gap. At sampling time, FSCG applies local k-NN feature correction and selects the best of multiple stochastic candidates according to their feature-space energy. In this way, the mask defines the anatomy, while FSCG steers samples toward the real US domain. Across three different datasets, FSCG reduces average FID64 by 56%, FID192 by 57%, and nearest-neighbour feature distance by 47% over standard conditional diffusion sampling, outperforming alternative inference-time guidance baselines. The results suggest that domain-aware feature representations can reveal and reduce realism gaps in medical diffusion synthesis without retraining the generator.  \nOur code is available at [https://github.com/marinadominguez/FSCG](https://github.com/marinadominguez/FSCG).  \nKeywords: Ultrasound · Synthetic Image Generation · Diffusion Models · Inference-time Guidance  \n1 Introduction  \nMedical ultrasound (US) is a widely used clinical imaging modality due to its low cost, portability, real-time acquisition, and absence of ionizing radiation [30] . However, developing robust machine-learning models for US remains challenging: image appearance varies substantially across operators, acquisition protocols, probes, anatomies, and clinical centres. Moreover, expert annotations are costly and data sharing is often limited by privacy constraints [29,25] . Realistic US synthesis could therefore support data augmentation, robustness analysis, and simulation when large curated datasets are not available [5,28 ,31] .  \n2 M. Domínguez et al.  \nMedical image synthesis, and ultrasound image generation in particular, has increasingly benefited from diffusion models [1, 10] . Most conditional US synthesis methods control the generated anatomy through masks, semantic maps, or related structural inputs [28,31] . These signals specify the spatial layout, but they do not explicitly enforce realistic B-mode appearance. Moreover, evaluating the realism of the images is challenging because common metrics, such as PSNR, SSIM and LPIPS, are sensitive to normalization, alignment, background regions, or blur [4], whereas distributional measures such as FID [7] rely on natural-image feature spaces that do not reflect US-specific appearance characteristics [14,23] . These limitations call for evaluation and guidance in representation spaces learned from the actual ultrasound data.  \nFigure 1 motivates our approach. We project real echocardiographic images [15] and samples from a conditioned diffusion transformer (DiT) into the feature space of OpenUS, a foundation model for US image analysis [34] . We use DiT as baseline model because it has shown excellent scalability and stateof-the-art image-generation performance on class-conditional benchmarks [21] . We consider a mask-conditioned DiT base","cbCaisMEmaqXuOQU","https://ap.wps.com/l/cbCaisMEmaqXuOQU","pdf",2370973,4,1,11,"English","en",105,"# Introduction\n## Problem: anatomy vs B-mode realism\n## Motivation from representation-space analysis\n# Method\n## Feature-Space Candidate Guidance (FSCG)\n## Feature bank scoring and candidate selection\n# Experiments\n## Datasets and evaluation metrics\n## Comparison with baseline guidance methods","[{\"question\":\"Why can conditional diffusion models produce anatomically plausible ultrasound images but still look unrealistic in B-mode?\",\"answer\":\"Anatomical plausibility focuses on structural alignment, while B-mode realism is governed by ultrasound-specific acquisition factors such as speckle texture, tissue contrast, and attenuation. Standard conditioning does not explicitly enforce these modality-dependent appearance traits.\"},{\"question\":\"What is Feature-Space Candidate Guidance (FSCG)?\",\"answer\":\"FSCG is a training-free inference-time sampling strategy that applies local k-NN feature correction in a frozen ultrasound feature space and selects the best among multiple stochastic candidates by feature-space energy. This steers samples toward the real ultrasound domain.\"},{\"question\":\"How does FSCG improve results compared with standard conditional diffusion guidance?\",\"answer\":\"Across three datasets, FSCG reduces average FID64 by 56% and FID192 by 57%, and lowers nearest-neighbour feature distance by 47% versus standard conditional diffusion sampling. It also outperforms alternative inference-time guidance baselines.\"}]",1784210114,28,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"feature-space-guided-diffusion-for-realistic-ultrasound-image-synthesis","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/feature-space-guided-diffusion-for-realistic-ultrasound-image-synthesis/86290/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why can conditional diffusion models produce anatomically plausible ultrasound images but still look unrealistic in B-mode?","Question",{"text":75,"@type":76},"Anatomical plausibility focuses on structural alignment, while B-mode realism is governed by ultrasound-specific acquisition factors such as speckle texture, tissue contrast, and attenuation. Standard conditioning does not explicitly enforce these modality-dependent appearance traits.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What is Feature-Space Candidate Guidance (FSCG)?",{"text":80,"@type":76},"FSCG is a training-free inference-time sampling strategy that applies local k-NN feature correction in a frozen ultrasound feature space and selects the best among multiple stochastic candidates by feature-space energy. This steers samples toward the real ultrasound domain.",{"name":82,"@type":73,"acceptedAnswer":83},"How does FSCG improve results compared with standard conditional diffusion guidance?",{"text":84,"@type":76},"Across three datasets, FSCG reduces average FID64 by 56% and FID192 by 57%, and lowers nearest-neighbour feature distance by 47% versus standard conditional diffusion sampling. It also outperforms alternative inference-time guidance baselines.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]