[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82386-en":3,"doc-seo-82386-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},82386,1099514068365,"Aurelia","https://ap-avatar.wpscdn.com/avatar/10000253d8d9f28188e?_k=1776742907772140068",8,"Research & Report","ALICE Learning a General-Purpose Pathology Foundation Model from Vision, Vision-Language, and Slide-Level Experts","Foundation models are reshaping computational pathology, yet their abilities are limited by pretraining objectives, data sources, and spatial scales, which fragments expertise across different backbones. ALICE is a unified foundation model trained via multi-stage agglomerative distillation that sequentially distills eight specialized teacher models across vision-only, vision–language, and slide-level modalities into one backbone. Pretraining uses 24,985,184 tile-level and 155,604 high-resolution pathology images, evaluated on 21 scenarios, 96 tasks, and 48 data sources, achieving the best average rank on matched tasks.","arXiv :2607 .09526v1 [ cs .CV] 10 Jul 2026  \nALICE: Learning a General-Purpose Pathology Foundation Model from Vision, Vision-Language, and Slide-Level Experts  \nJiawen Li 1∗, Tian Guan 1∗, Huijuan Shi3∗, Xitong Ling 1 , Mingxi Fu 1 , Anjia Han3+ , Chao He2+ , Yonghong He 1 ,4+  \n1Institute of Biopharmaceutical and Health Engineering, Tsinghua Shenzhen International Graduate School, Tsinghua University, Shenzhen, China  \n2Department of Engineering Science, University of Oxford, Oxford, UK  \n3Department of Pathology, The First Affiliated Hospital of Sun Yat-sen University, Guangzhou, China  \n4Medical Optical Technology R&D Center, Research Institute of Tsinghua, Pearl River Delta, Guangzhou, China  \n∗ Contributed equally  \n+ Corresponding Authors:  \nAnjia Han ([hananjia@mail.sysu.edu.cn](hananjia@mail.sysu.edu.cn)), Chao He ([chao.he@eng.ox.ac.uk](chao.he@eng.ox.ac.uk))  \nYonghong He ([heyh@sz.tsinghua.edu.cn](heyh@sz.tsinghua.edu.cn))  \nAbstract  \nFoundation models are reshaping computational pathology, yet their capabilities remain shaped by pretraining objectives, data sources, and spatial scales, fragmenting complementary expertise across separate backbones. Here we present ALICE, a unified foundation model trained through multi-stage agglomerative distillation that sequentially distills eight vision-only, vision–language, and slide-level teacher models into dedicated modules of a single backbone. ALICE is pretrained on 24,985,184 tile-level pathology images and 155,604 high-resolution images, and evaluated across 21 task scenarios, 96 downstream tasks, and 48 data sources, spanning region-of-interest tissue analysis, vision–language multimodal evaluation, and whole-slide clinical assessment. In all three evaluation settings, ALICE achieved the best average rank among task-matched pathology foundation models. These results demonstrate that agglomerative distillation can consolidate complementary capabilities from specialized models into a unified backbone for broad computational pathology applications. The model is available at [https://github.com/WonderLandxD/ALICE](https://github.com/WonderLandxD/ALICE).  \nIntroduction  \nHistopathological assessment of tissue remains the cornerstone of clinical oncology, forming the basis for diagnostic, prognostic, and treatment decisions across many cancer types 1. Digitizing slides into whole-slide images (WSIs) has created unprecedented opportunities to develop computational tools that enhance pathology analysis at scale 2,3 . Prior deep learning methods have shown potential for specific tasks such as tumor detection, molecular biomarker prediction, and survival analysis 4–7. However, most of these models are trained from scratch on task-specific, typically small labeled cohorts, limiting their performance and generalization capabilities. Largescale foundation models pretrained on massive histopathology datasets through self-supervised or multimodal learning have fundamentally changed this paradigm, providing transferable vision representations adaptable to a wide range of downstream clinical tasks with substantially reduced annotation requirements 8–11. Consequently, a growing number of pathology foundation models (PFMs) have been developed, demonstrating substantial potential across diverse clinical applications.  \nDespite their common goal of learning general-purpose representations, existing PFMs are developed under different pretraining paradigms, leading to distinct but fragmented capabilities 12,13 . For instance, vision-only models, such as UNI 8 and Virchow 14 , rely on self-supervised learning to capture dense morphological representations, but they lack explicit alignment with pathological concepts and language-level semantics. Vision-language models, such as CONCH 9 and MUSK 15 , rely on multimodal contrastive or generative learning to align visual features with textual semantics, but they underperform on fine-grained visual discrimination tasks 16. Slide-level models, suc","cbCaij25QOshq30p","https://ap.wps.com/l/cbCaij25QOshq30p","pdf",5294118,2,1,89,"English","en",105,"# Abstract\n# Introduction\n## Existing pathology foundation models and their gaps\n## Agglomerative distillation as a solution\n# ALICE overview and training objective","[{\"question\":\"What problem does ALICE address in computational pathology foundation models?\",\"answer\":\"Existing pathology foundation models are trained under different paradigms, producing fragmented capabilities across local morphology, language-aligned concepts, and whole-slide context. ALICE aims to consolidate these complementary abilities into one unified backbone.\"},{\"question\":\"How is ALICE trained to unify multiple expert models?\",\"answer\":\"ALICE uses multi-stage agglomerative distillation, sequentially distilling eight teacher models from vision-only, vision–language, and slide-level settings into dedicated modules within a single backbone.\"},{\"question\":\"What datasets and evaluation scope are used to validate ALICE?\",\"answer\":\"ALICE is pretrained on 24,985,184 tile-level pathology images and 155,604 high-resolution images, then evaluated across 21 task scenarios, 96 downstream tasks, and 48 data sources covering ROI analysis, multimodal evaluation, and whole-slide clinical assessment.\"}]",1784180075,224,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"alice-learning-a-general-purpose-pathology-foundation-model-from-vision-vision-language-and-slide-level-experts","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/alice-learning-a-general-purpose-pathology-foundation-model-from-vision-vision-language-and-slide-level-experts/82386/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-22","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does ALICE address in computational pathology foundation models?","Question",{"text":75,"@type":76},"Existing pathology foundation models are trained under different paradigms, producing fragmented capabilities across local morphology, language-aligned concepts, and whole-slide context. ALICE aims to consolidate these complementary abilities into one unified backbone.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How is ALICE trained to unify multiple expert models?",{"text":80,"@type":76},"ALICE uses multi-stage agglomerative distillation, sequentially distilling eight teacher models from vision-only, vision–language, and slide-level settings into dedicated modules within a single backbone.",{"name":82,"@type":73,"acceptedAnswer":83},"What datasets and evaluation scope are used to validate ALICE?",{"text":84,"@type":76},"ALICE is pretrained on 24,985,184 tile-level pathology images and 155,604 high-resolution images, then evaluated across 21 task scenarios, 96 downstream tasks, and 48 data sources covering ROI analysis, multimodal evaluation, and whole-slide clinical assessment.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]