[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85892-en":3,"doc-seo-85892-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85892,687197207639,"Asher","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift","Foundation models increasingly serve as image feature extractors in mammography, yet their robustness under external domain shift is not well established. This work benchmarks 15 foundation-model backbones across breast density, BI-RADS severity, and cancer status using a unified frozen-backbone linear-probe protocol. Training uses three source datasets; evaluation uses 12 compatible OOD datasets after label harmonization. Mammo-specific vision-language models deliver the strongest mean OOD performance, though robustness is not explained by mammography exposure alone. Dataset-level OOD analysis and feature-space inspection further reveal heterogeneous behavior across datasets, supporting OOD dataset evaluation as a central criterion. Code is publicly available.","arXiv :2607 . 10358v1 [ cs .CV] 11 Jul 2026  \nBenchmarking the Robustness of Foundation Models for Mammography under Domain Shift  \nGiang Nguyen*, \\#, 1, Raghav Mehta *, \\#, 2, Emma A.M. Stanley2 , Tian Xia2 , Thi Hao Nguyen3 , Hieu Pham 1,4,5 , and Ben Glocker2  \n1 College of Engineering and Computer Science, VinUniversity, Hanoi, Vietnam  \n2 Imperial College London, London, UK  \n3 Radiology Department, Vietnam National Cancer Hospital, Hanoi, Vietnam  \n4VinUni-Illinois Smart Health Center, VinUniversity, Hanoi, Vietnam  \n5 The Computer Vision and Medical AI Lab, VinUniversity, Hanoi, Vietnam  \n*  \nEqual contribution  \n\\# Corresponding authors: [23giang.ns@vinuni.edu.vn](23giang.ns@vinuni.edu.vn), [raghav.mehta@imperial.ac.uk](raghav.mehta@imperial.ac.uk)  \nAbstract  \nFoundation models are increasingly used as image feature extractors for mammography, but their robustness under external domain shift remains unclear. We benchmark 15 foundation-model backbones across breast density, BI-RADS severity, and cancer status using a unified frozen-backbone linear-probe protocol, training on  \n3 source datasets and evaluating on 12 task-compatible out-of-distribution (OOD) datasets after label harmonization. Mammography-specific vision-language models (Mammo-FM and MaMA) provide the strongest mean OOD performance, but robustness is not explained by mammography exposure alone. DINOv3 remains a competitive vision-only baseline, and mammography-adapted pretraining does not consistently improve generalization. Dataset-level analysis further shows that even leading models show heterogeneous performance across datasets. Feature-space inspection reveals that useful representations can preserve clinical signal while retaining dataset and acquisition structure. These findings highlight dataset-level OOD evaluation asa central criterion for assessing mammography representations. Our code is publicly available: [https://github.com/biomedia-mira/mammo-ood](https://github.com/biomedia-mira/mammo-ood).  \nKeywords: Mammography · Foundation models · Out-of-distribution generalization · Domain shift · Vision-language models  \n1 Introduction  \nFoundation models (FMs) offer a practical route towards broadly applicable mammography representations. A typical workflow freezes a pretrained backbone, trains a lightweight classification head on a labeled source dataset, and applies the classifier to external data. A single backbone may support multiple downstream mammography tasks while reducing  \nFigure 1: Overview of the mammography foundation-model benchmark.  \ntask-specific architectural design. However, the utility of an FM depends on whether its pretrained representation remains useful under distribution shift.  \nPrior work [4 , 9 , 12 , 23 , 29] has shown that large-scale or modality-specific pretraining does not by itself guarantee robust out-of-distribution (OOD) generalization. This issue is amplified in mammography by heterogeneous clinical labels: density, BI-RADS, and cancer status differ in granularity, prevalence, and annotation protocols. Thus, indistribution (ID) performance may not translate to OOD generalization. To evaluate whether mammography representations generalize beyond their source data, we establish a benchmark for FM encoders across heterogeneous clinical tasks. We compare natural-image self-supervised learning (SSL) [27 , 33], general radiology SSL [25 , 30], mammography adapted SSL [14 , 27 , 33 , 15], general medical vision-language models (VLM) [19 , 32 , 35], and mammography specific VLM [7 , 8 , 10 , 11] under frozen-backbone linear-probe protocol. Our contributions are threefold:  \n• We construct a harmonized benchmark from 15 public mammography datasets spanning 12 countries/regions. It covers three different downstream tasks: imagelevel density, exam-level BI-RADS, and exam-level cancer status.  \n• We evaluate 15 foundation-model backbones under a unified frozen-backbone sourceto-external ID/OOD protocol across these thre","cbCaivn4iO24JkgC","https://ap.wps.com/l/cbCaivn4iO24JkgC","pdf",4798724,2,1,11,"English","en",105,"# 1 Introduction\n# 2 Benchmark Setup\n## 2.1 Benchmark Datasets and Tasks","[{\"question\":\"How does the benchmark evaluate foundation models for mammography under domain shift?\",\"answer\":\"It uses a unified frozen-backbone linear-probe protocol. Models are trained on three source datasets and evaluated on 12 task-compatible out-of-distribution datasets after label harmonization.\"},{\"question\":\"Which model types show the best out-of-distribution performance?\",\"answer\":\"Mammography-specific vision-language models (Mammo-FM and MaMA) achieve the strongest mean OOD performance. DINOv3 remains a competitive vision-only baseline.\"},{\"question\":\"Why is mammography exposure alone insufficient to explain robustness?\",\"answer\":\"The study finds that robustness is influenced by factors beyond mammography exposure, including the pretraining objective, whether the model is vision-only vs. vision-language, and the diversity of the pretraining dataset.\"}]",1784206991,28,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"benchmarking-the-robustness-of-foundation-models-for-mammography-under-domain-shift","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/benchmarking-the-robustness-of-foundation-models-for-mammography-under-domain-shift/85892/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"How does the benchmark evaluate foundation models for mammography under domain shift?","Question",{"text":75,"@type":76},"It uses a unified frozen-backbone linear-probe protocol. Models are trained on three source datasets and evaluated on 12 task-compatible out-of-distribution datasets after label harmonization.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Which model types show the best out-of-distribution performance?",{"text":80,"@type":76},"Mammography-specific vision-language models (Mammo-FM and MaMA) achieve the strongest mean OOD performance. DINOv3 remains a competitive vision-only baseline.",{"name":82,"@type":73,"acceptedAnswer":83},"Why is mammography exposure alone insufficient to explain robustness?",{"text":84,"@type":76},"The study finds that robustness is influenced by factors beyond mammography exposure, including the pretraining objective, whether the model is vision-only vs. vision-language, and the diversity of the pretraining dataset.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]