[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-81990-en":3,"doc-seo-81990-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},81990,687197207639,"Asher","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","Vision Foundation Models in Radiology: A Scoping Review of Data, Methodology, Evaluation and Clinical Translation","Vision foundation models (VFMs) are increasingly being developed for radiological imaging, yet their definition, development and evaluation remain heterogeneous. A PRISMA-ScR scoping review included 67 peer-reviewed studies (Jan 2017–Mar 2026) of models trained exclusively on radiology imaging data. Evidence was mapped across data scale/heterogeneity, architectural and pretraining scalability, and downstream transferability. Results highlight mainly brain MRI, thoracoabdominal CT, and chest X-ray; evaluation favors segmentation and classification, while cross-center, cross-scanner, anatomical and modality-shift validation is inconsistently reported. Clinical translation is limited by representativeness, incomplete reporting, benchmark heterogeneity, and deployment-oriented evaluation gaps.","arXiv :2607 .072 19v 1 [ cs .CV] 8 Jul 2026  \nVision Foundation Models in Radiology: A Scoping Review of Data, Methodology, Evaluation and Clinical Translation  \nAlejandro Vergara-Richart 1,2, ∗ , Xavier Rafael-Palou1, ∗ , Almudena Fuster-Matanzo1 ,  \nIgnacio Iborra Roncales 1 , Ángel Alberich-Bayarri 1 , Ana Jiménez-Pastor 1  \n1 Quantitative Imaging Biomarkers in Medicine (Quibim S.L.), Valencia, Spain  \n2Universitat Politècnica de València, Valencia, Spain  \n∗ These authors contributed equally to this work.  \nCorrespondence: Alejandro Vergara-Richart, [alejandrovergara@quibim.com](alejandrovergara@quibim.com)  \nKeywords: foundation models; radiology; review; medical imaging  \nAbstract  \nVision foundation models (VFMs) are increasingly being developed for radiological imaging, yet their definition, development and evaluation remain heterogeneous. We conducted a PRISMAScR scoping review of peer-reviewed studies published between January 2017 and March 2026 describing foundation models trained exclusively on radiological imaging data. Sixty-seven studies were included and mapped across three pillars: data scale and heterogeneity, architectural and pretraining scalability, and downstream transferability and generalization. Datasets primarily covered brain MRI, thoracoabdominal CT, and chest X-ray, ranging from fewer than 100,000 samples to multi-million-image cohorts. Transformer-based architectures and self-supervised pretraining predominated, particularly masked image modeling, contrastive learning and multi-stage approaches. Evaluation focused mainly on segmentation and classification, whereas cross-center, cross-scanner, anatomical and modality-shift validation was inconsistently reported. Alignment with FUTURE-AI principles was uneven. Overall, radiology-specific VFMs show promising transferability, but clinical translation remains constrained by limited data representativeness, heterogeneous benchmarks, incomplete reporting and insufficient deployment-oriented evaluation.  \n1 Introduction  \nOriginating in natural language processing and popularized through large language models [1], foundation models (FMs) represent a shift toward learning general-purpose representations from large-scale data. Unlike task-specific models optimized for a single objective, FMs are trained on diverse and extensive datasets with the objective of capturing transferable features that can be adapted to multiple downstream applications.  \nThis paradigm shift has been extended to computer vision through vision foundation models (VFMs), which are high-capacity neural networks trained on large-scale, heterogeneous, and often unlabeled imaging data to learn reusable representations that can be efficiently adapted to diverse specific tasks [2] . These properties are particularly relevant in radiology imaging analysis, where data scarcity, high annotation  \ncosts, and substantial variability across institutions, scanners and acquisition protocols hinder predictive model generalization and clinical translation [3] . Furthermore, such variability induces distributional shifts that significantly degrade performance under external validation and real-world deployment settings.  \nSelf-supervised learning (SSL) has become the dominant strategy for pretraining models in medical imaging due to its ability to exploit the intrinsic structure of unlabeled data [4] . In medical imaging, SSL approaches are typically grouped into two broad families: (i) reconstruction-based methods, such as masked image modeling and related generative objectives [5], and (ii) contrastive learning frameworks that promote invariance across augmented views, modalities, or patient-level representations [6] . Complementary, weakly supervised learning strategies, which exploit coarse or inexact labels, have also contributed to scaling representation learning in radiology [7] .  \nBuilt upon these pretraining strategies, VFMs are typically implemented using large-capacity architectures (ofte","cbCaibu0VHdUYbL5","https://ap.wps.com/l/cbCaibu0VHdUYbL5","pdf",1998218,7,1,33,"English","en",105,"# Abstract\n# Introduction","[{\"question\":\"What is the goal of the scoping review on vision foundation models in radiology?\",\"answer\":\"The review maps how radiology-specific vision foundation models are defined, developed, and evaluated by synthesizing peer-reviewed evidence across data, methodology, evaluation, and clinical translation aspects.\"},{\"question\":\"Which types of radiology data and datasets are most commonly used in the included studies?\",\"answer\":\"The reviewed studies mainly cover brain MRI, thoracoabdominal CT, and chest X-ray, with dataset sizes ranging from fewer than 100,000 samples to multi-million-image cohorts.\"},{\"question\":\"How are these models evaluated, and what validation gaps are reported?\",\"answer\":\"Evaluation mainly focuses on segmentation and classification. Cross-center, cross-scanner, anatomical, and modality-shift validation is inconsistently reported, and deployment-oriented evaluation remains insufficient.\"}]",1784177450,83,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"vision-foundation-models-in-radiology-a-scoping-review-of-data-methodology-evaluation-and-clinical-translation","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/vision-foundation-models-in-radiology-a-scoping-review-of-data-methodology-evaluation-and-clinical-translation/81990/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-03","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What is the goal of the scoping review on vision foundation models in radiology?","Question",{"text":76,"@type":77},"The review maps how radiology-specific vision foundation models are defined, developed, and evaluated by synthesizing peer-reviewed evidence across data, methodology, evaluation, and clinical translation aspects.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"Which types of radiology data and datasets are most commonly used in the included studies?",{"text":81,"@type":77},"The reviewed studies mainly cover brain MRI, thoracoabdominal CT, and chest X-ray, with dataset sizes ranging from fewer than 100,000 samples to multi-million-image cohorts.",{"name":83,"@type":74,"acceptedAnswer":84},"How are these models evaluated, and what validation gaps are reported?",{"text":85,"@type":77},"Evaluation mainly focuses on segmentation and classification. Cross-center, cross-scanner, anatomical, and modality-shift validation is inconsistently reported, and deployment-oriented evaluation remains insufficient.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":107,"slug":138},19,"General","general"]