[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-122139-en":3,"doc-seo-122139-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},122139,687197207639,"Asher","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",6,"Technology","Masked Capsule Autoencoders","Masked Capsule Autoencoders (MCAE) introduce the first Capsule Network built on modern self-supervised pretraining in the masked image modelling (MIM) paradigm. By reformulating Capsule Networks to include masked image modelling before supervised fine-tuning, the approach improves scalability to more complex, realistic-resolution data. Experiments and ablations show Capsule Networks benefit from self-supervised pretraining similarly to CNNs and Vision Transformers. Pretraining on Imagenette yields state-of-the-art Capsule results with a reported 9% gain over a baseline model.","Masked Capsule Autoencoders  \nMiles Everett [miles. everett@abdn. ac.uk](miles. everett@abdn. ac.uk)  \nDepartment of Computing Science University of Aberdeen, UK  \nMingjun Zhong [mingjun.zhong@abdn. ac.uk](mingjun.zhong@abdn. ac.uk)  \nDepartment of Computing Science University of Aberdeen, UK  \nGeorgios Leontidis [georgios.leontidis@abdn. ac.uk](georgios.leontidis@abdn. ac.uk)  \nInterdisciplinary Institute Department of Computing Science University of Aberdeen, UK  \nReviewed on OpenReview: [https: // openreview. net/ forum? id= JHxrh00W1j](https: // openreview. net/ forum? id= JHxrh00W1j)  \nAbstract  \nWe propose Masked Capsule Autoencoders (MCAE), the first Capsule Network that utilisespretraining in a modern self-supervised paradigm, specifically the masked image modelling framework. Capsule Networks have emerged as a powerful alternative to Convolutional Neural Networks (CNNs) . They have shown favourable properties when compared to Vision Transformers (ViT), but have struggled to effectively learn when presented with more complex data. This has led to Capsule Network models that do not scale to modern tasks.  \nOur proposed MCAE model alleviates this issue by reformulating the Capsule Network to use masked image modelling as a pretraining stage before finetuning in a supervised manner.  \nAcross several experiments and ablations studies we demonstrate that similarly to CNNsand ViTs, Capsule Networks can also benefit from self-supervised pretraining, paving the way for further advancements in this neural network domain. For instance, by pretraining on the Imagenette dataset—consisting of 10 classes of Imagenet-sized images—we achieve state-of-the-art results for Capsule Networks, demonstrating a 9% improvement compared to our baseline model. Thus, we propose that Capsule Networks benefit from and should be trained within a masked image modelling framework, using a novel capsule decoder, to enhance a Capsule Network’s performance on realistically sized images.  \n1 Introduction  \nCapsule Networks are an evolution of Convolutional Neural Networks (CNNs) which remove pooling operations and replace scalar neurons with a fixed number of vector or matrix representations known as capsules at each location in the feature map. At each location, there will be multiple capsules, each theoretically representing a different concept. Each of these capsules has a corresponding activation value between 0 and 1 . The activation value measures the probability (from 0 to 1) that the concept represented by the capsule exists. Capsule Networks have shown promising signs, such as being naturally strong in invariant and equivariant tasks (Sabour et al., 2017; Hinton et al., 2018; De Sousa Ribeiro et al., 2020; Hahn et al. , 2019; Ribeiro et al., 2020; 2022) while having low parameter counts, but have yet to scale to more complex datasets with realistic resolutions that CNNs and Vision Transformers (ViTs) are typically benchmarked on.  \nMasked Image Modelling (MIM) is a Self Supervised Learning (SSL) technique with roots in language modelling (Devlin et al., 2018) . In language modelling, words are removed from passages of text, the network is then trained to predict the correct words to fill in the gaps. This technique can be extended to  \nSplit Image to Patches  \nFlatten  \nRandom Patch Selection  \nPer Patch Convolutional Feature Map  \nPer Patch Capsule Representations  \nMask tokens reinserted  \nReshape  \nFigure 1: Our Masked Capsule Autoencoder architecture. During pretraining we randomly select a number of patches from the original image to be processed. The Capsule Network will then create a representation for each patch. Masked patch capsule representations are then re-added before the capsule decoder, where the unmasked capsules can contribute to the masked positions, which are finally decoded by a single linear layer to the original patch dimensions. The pretraining objective is the mean squared error between thereconstructed patches and the ta","cbCaisDy0zEHmFiz","https://ap.wps.com/l/cbCaisDy0zEHmFiz","pdf",849764,1,15,"English","en",105,"# Abstract\n# 1 Introduction\n## Capsule Networks and limitations on complex data\n## Masked Image Modelling in self-supervised learning\n## Proposed MCAE architecture and pretraining objective","[{\"question\":\"What are Masked Capsule Autoencoders (MCAE)?\",\"answer\":\"Masked Capsule Autoencoders are a Capsule Network design that incorporates masked image modelling as a self-supervised pretraining stage before supervised fine-tuning.\"},{\"question\":\"How does MCAE address Capsule Networks struggling with complex data?\",\"answer\":\"It improves training by reformulating Capsule Networks to learn richer region-level representations via masked image modelling prior to activating global class capsules during fine-tuning.\"},{\"question\":\"What performance gains does the paper report using pretraining?\",\"answer\":\"Pretraining on the Imagenette dataset enables state-of-the-art results for Capsule Networks, with a reported 9% improvement compared to the baseline model.\"}]","Masked Capsule Autoencoders | PDF",1785809000,38,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"masked-capsule-autoencoders","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/technology/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/masked-capsule-autoencoders/122139/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What are Masked Capsule Autoencoders (MCAE)?","Question",{"text":75,"@type":76},"Masked Capsule Autoencoders are a Capsule Network design that incorporates masked image modelling as a self-supervised pretraining stage before supervised fine-tuning.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does MCAE address Capsule Networks struggling with complex data?",{"text":80,"@type":76},"It improves training by reformulating Capsule Networks to learn richer region-level representations via masked image modelling prior to activating global class capsules during fine-tuning.",{"name":82,"@type":73,"acceptedAnswer":83},"What performance gains does the paper report using pretraining?",{"text":84,"@type":76},"Pretraining on the Imagenette dataset enables state-of-the-art results for Capsule Networks, with a reported 9% improvement compared to the baseline model.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,113,118,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":111,"slug":112},50,"technology",{"id":114,"doc_module":4,"doc_module_name":46,"category_name":115,"show_sort_weight":116,"slug":117},7,"Healthcare",40,"healthcare",{"id":119,"doc_module":4,"doc_module_name":46,"category_name":120,"show_sort_weight":121,"slug":122},8,"Research & Report",30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]