[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-81613-en":3,"doc-seo-81613-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},81613,34359740700684,"Finn","https://ap-avatar.wpscdn.com/avatar/1f400023980c374ae676?_k=1777273430885731487",8,"Research & Report","LUMOS Latent Universal Medical Priors for Segmentation","General vision foundation models (VFMs) are widely trained on natural images, so medical image segmentation is often viewed as requiring costly adaptation or domain-specific fine-tuning. This work challenges that assumption by testing whether low-level priors for anatomical delineation are already embedded in frozen VFMs. The study finds transferable visual regularities without medical supervision and introduces LUMOS, which amplifies these priors for conventional medical segmentors using Pathfinder and Inspiror. Experiments across medical datasets show robust, token-based spatial guidance from frozen token spaces.","LUMOS: Latent Universal Medical Priors for  \nSegmentation  \nZhuonan Liang 1 , Wei Guo 1 , Jie Gan 1 , Yaxuan Song 1 , Runnan Chen 1 , Hang Chang2 ,3 , and Weidong Cai 1  \narXiv :2603 .01115v2 [ cs .CV] 10 Jul 2026  \nAbstract—General vision foundation models (VFMs) have been primarily developed on natural images, and their utility for medical image segmentation is therefore often considered to depend on costly adaptation or domain-specific fine-tuning. In this paper, we revisit this assumption from a different perspective: rather than requiring VFM segmentors to relearn visual regularities, we investigate whether the low-level visual priors necessary for anatomical delineation already lie dormant within general VFMs. We observe that frozen VFMs, despite lacking medical supervision, encode transferable visual regularities. These properties are not exclusive to natural images but are also fundamental to medical image understanding. Motivated by this observation, we propose Latent Universal Medical PriOrs for Segmentation (LUMOS), a novel framework that amplifies general VFM priors to conventional medical segmentors. LUMOS consists of two key components: (1) Pathfinder that distills visual cues from a frozen vision foundation model, and (2) Inspiror that sparks the conventional medical networks with spatial guidance from distilled visual regularities. In this way, the segmentor is relieved from learning complex visual regularities entirely from limited medical annotations and can instead focus on taskspecific anatomical delineation. Across diverse medical datasets and token-based VFMs, LUMOS shows that general VFMscan serve as spatial prior generators when their frozen token spaces preserve patch-level pattern relevance. DINO provides stable matched-backbone gains, while SigLIP exposes VFMspecific sensitivity caused by its different token granularity and representation objective.  \nIndex Terms—Medical image segmentation, vision foundation model, spatial guidance, DINOv3, weak annotation.  \nI. INTRODUCTION  \nVision foundation models (VFMs) have reshaped computer vision through general representations learned at scale [1]–[4] . Token-based VFMs, such as DINOv3 and SigLIP, provide dense features with global and local cues that may benefit medical segmentation [1], [2], [5], [6] . Yet their token semantics are not directly aligned with medical tasks [4], [5], [7], and fine-tuning them can require compute and annotations that remain scarce in medical imaging [8] . Conventional medical segmentors, including convolutional and transformer variants, remain strong at modality-specific feature learning [9]–[14] . The question is therefore how to use VFM information without forcing a medical segmentor to absorb mismatched semantics. DINO-style ViTs provide a clue: their self-attention exposes object layout and boundary information, and their patch  \n1The University of Sydney, Sydney, NSW 2006, Australia.  \n2Biological System & Engineering Division, Lawrence Berkeley National Laboratory, Berkeley, CA 94720, USA.  \n3Berkeley Biomedical Data Science Center, Lawrence Berkeley National Laboratory, Berkeley, CA 94720, USA.  \nCorresponding author: Weidong Cai ([tom.cai@sydney.edu.au](tom.cai@sydney.edu.au)).  \nfeatures serve as dense descriptors for co-segmentation and correspondence [15], [16] . Recent medical studies likewise suggest that frozen DINO features preserve structural cues useful for lightweight segmentation readout [4], [17] . We therefore ask how the useful visual priors in native VFMscan be extracted and injected into medical segmentors in a controlled way.  \nBy leveraging the rich token representations from foundation models, we can efficiently guide the segmentation process while preserving the inductive biases of medical image segmentation architectures. The insights of this approach are twofold: (1) Foundation models can serve as visualprior extractors that provide spatial guidance instead of direct medical semantics; (2) The ded","cbCaieOXHh4EqtTX","https://ap.wps.com/l/cbCaieOXHh4EqtTX","pdf",32194828,4,1,14,"English","en",105,"# Introduction\n## Motivation and Problem Setting\n## Proposed Framework: LUMOS\n## Contributions Overview","[{\"question\":\"Why are medical segmentation models often thought to need expensive VFM adaptation or fine-tuning?\",\"answer\":\"Because most VFMs are trained on natural images, their token semantics are not directly aligned with medical tasks, and fine-tuning may require compute and annotations that are scarce in medical imaging.\"},{\"question\":\"What does the paper claim about frozen vision foundation models for medical segmentation?\",\"answer\":\"Frozen VFMs, despite lacking medical supervision, encode transferable visual regularities that are useful beyond natural images, supporting medical image understanding and anatomical delineation.\"},{\"question\":\"How does LUMOS inject general VFM priors into conventional medical segmentors?\",\"answer\":\"LUMOS uses Pathfinder to distill visual cues from a frozen VFM and place them onto a spatial guide mask via a regularity-prototype mechanism, then Inspiror gates medical network feature activations using this spatial mask to inject priors while preserving segmentation inductive biases.\"}]",1784174786,35,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"lumos-latent-universal-medical-priors-for-segmentation","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/lumos-latent-universal-medical-priors-for-segmentation/81613/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-26","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why are medical segmentation models often thought to need expensive VFM adaptation or fine-tuning?","Question",{"text":75,"@type":76},"Because most VFMs are trained on natural images, their token semantics are not directly aligned with medical tasks, and fine-tuning may require compute and annotations that are scarce in medical imaging.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What does the paper claim about frozen vision foundation models for medical segmentation?",{"text":80,"@type":76},"Frozen VFMs, despite lacking medical supervision, encode transferable visual regularities that are useful beyond natural images, supporting medical image understanding and anatomical delineation.",{"name":82,"@type":73,"acceptedAnswer":83},"How does LUMOS inject general VFM priors into conventional medical segmentors?",{"text":84,"@type":76},"LUMOS uses Pathfinder to distill visual cues from a frozen VFM and place them onto a spatial guide mask via a regularity-prototype mechanism, then Inspiror gates medical network feature activations using this spatial mask to inject priors while preserving segmentation inductive biases.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]