[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-81685-en":3,"doc-seo-81685-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},81685,1649267921044,"Ava Thompson","https://us-avatar.wpscdn.com/avatar/1800007509477c92dfb?_k=1782875107921204101",8,"Research & Report","3D Masked Autoencoders are Robust Learners of Volumetric and Multimodal Cellular Representations for Microscopy","Self-supervised learning in fluorescence microscopy often compresses inherently three-dimensional cells into 2D projections. This work presents a systematic comparison of 2D and 3D masked autoencoders (MAE-2D vs. MAE-3D) on volumetric microscopy data with matched training settings. MAE-3D improves downstream single-cell tasks over 2D max-projection and slice-based baselines. Aligning image features with a pretrained protein language model enables stronger cross-modal gains for volumetric models. Channel cross-attention and frequency-domain regularization are key to exploiting 3D spatial context.","arXiv :2606 .23964v2 [ cs .LG] 9 Jul 2026  \n3D Masked Autoencoders are Robust Learners of Volumetric and Multimodal Cellular Representations for Microscopy  \nAmirhossein Kardoost 1 ,5 , Lion Gleiter 1 , Tingying Peng 1 , and  \nCarsten Marr 1 ,2 ,3 ,4 ,5  \n1 Institute of AI for Health & Helmholtz AI, Computational Health Center,  \nHelmholtz Munich – German Research Center for Environmental Health, Neuherberg, Germany  \n2 Department of Medicine III, Ludwig-Maximilian-University Hospital, Munich, Germany  \n3 Department of Physics, Ludwig-Maximilian-University, Munich, Germany  \n4 German Cancer Consortium (DKTK), partner site Munich, Germany  \n5 Munich Center for Machine Learning (MCML), Munich, Germany  \nAbstract. Self-supervised learning in fluorescence microscopy often relies on 2D projections, despite the inherently three-dimensional nature of cells. We present a systematic comparison of 2D and 3D masked autoencoders (MAE-2D vs. MAE-3D) on volumetric microscopy data. Under matched architectures and training protocols, MAE-3D consistently outperforms 2D max-projection and slice-based variants on downstream single-cell tasks. We further align visual representations with a pretrained protein language model (ESM2) and show that cross-modal supervision yields larger gains for volumetric models. Channel cross-attention and frequency-domain regularization are critical for leveraging 3D spatial context. On protein–protein interaction prediction, our best model achieves a ROC–AUC of 0.86, while on protein localization it reaches an AUCmicro of 0.95 and an F1micro of 0.74, demonstrating competitive performance on both tasks. Overall, our findings highlight the potential of volumetric modeling and multimodal alignment for representation learning in single-cell microscopy.  \nKeywords: Volumetric representation learning · Multimodal learning · Single-cell microscopy  \nCorresponding authors:  \n{amirhossein.kardoost, tingying.peng, [carsten.marr}@helmholtz-munich.de](carsten.marr}@helmholtz-munich.de)  \n2 A. Kardoost et al.  \n1 Introduction  \nCells constitute the fundamental building blocks of tissues and organs. Their function is tightly linked to subcellular structure and spatial organization. Understanding cellular organization remains a central challenge in biology. Despite extensive studies [13, 15 ,7 ,9 , 1], deciphering subcellular architecture and protein localization remains complex, particularly in high-dimensional imaging data. Fluorescence microscopy [5,3 ,20] enables visualization of intracellular structures by tagging proteins and organelles with fluorescent markers. Largescale resources such as JUMP [3], OpenCell [5], WTC-11 [20], and the Human Protein Atlas (HPA) [18] provide multi-channel imaging data capturing rich subcellular organization. Notably, OpenCell and WTC-11 consist of volumetric z-stacks. However, many representation learning approaches such as Subcell [9] and DINO4Cell [7] operate on 2D projections of these volumes, discarding depth-resolved structural information. We investigate the role of volumetric modeling for the learning of cellular representations. On OpenCell [5], we systematically compare 2D and 3D masked autoencoder (MAE) [22, 13] models. We demonstrate that preserving full 3D structure yields more informative representations and consistently improves downstream performance compared to 2D max-projection and even slice-based inputs. Beyond purely visual modeling, we explore multimodal integration by incorporating protein sequence information via a pretrained protein language model (PLM) such as ESM2 [14] . By aligning image features with protein embeddings, we infuse the representation space with biologically grounded structural priors. We demonstrate that sequence-level supervision enhances representation quality, particularly when coupled with volumetric modeling.  \nOur contributions are threefold: (1) We demonstrate that 3D MAE models outperform 2D counterparts across two downstream tasks. (2) We s","cbCainEAErveEedr","https://ap.wps.com/l/cbCainEAErveEedr","pdf",2139779,2,1,11,"English","en",105,"# Introduction\n# Related Work","[{\"question\":\"What problem does the document address in fluorescence microscopy representation learning?\",\"answer\":\"Fluorescence microscopy cells are inherently three-dimensional, but many self-supervised approaches rely on 2D projections, which discard depth-resolved structural information. The document studies how to better learn cellular representations from volumetric data.\"},{\"question\":\"How do 3D masked autoencoders compare to 2D variants in downstream performance?\",\"answer\":\"With matched architectures and training protocols, MAE-3D consistently outperforms 2D max-projection and slice-based variants on downstream single-cell tasks.\"},{\"question\":\"How does multimodal alignment with a protein language model improve learning?\",\"answer\":\"By aligning visual representations with embeddings from a pretrained protein language model (e.g., ESM2), cross-modal supervision yields larger gains for volumetric models. The document highlights that representation guidance and reconstruction benefits improve results.\"}]",1784175406,28,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"3d-masked-autoencoders-are-robust-learners-of-volumetric-and-multimodal-cellular-representations-for-microscopy","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/3d-masked-autoencoders-are-robust-learners-of-volumetric-and-multimodal-cellular-representations-for-microscopy/81685/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-23","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the document address in fluorescence microscopy representation learning?","Question",{"text":75,"@type":76},"Fluorescence microscopy cells are inherently three-dimensional, but many self-supervised approaches rely on 2D projections, which discard depth-resolved structural information. The document studies how to better learn cellular representations from volumetric data.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How do 3D masked autoencoders compare to 2D variants in downstream performance?",{"text":80,"@type":76},"With matched architectures and training protocols, MAE-3D consistently outperforms 2D max-projection and slice-based variants on downstream single-cell tasks.",{"name":82,"@type":73,"acceptedAnswer":83},"How does multimodal alignment with a protein language model improve learning?",{"text":84,"@type":76},"By aligning visual representations with embeddings from a pretrained protein language model (e.g., ESM2), cross-modal supervision yields larger gains for volumetric models. The document highlights that representation guidance and reconstruction benefits improve results.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]