[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82396-en":3,"doc-seo-82396-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},82396,1099514068365,"Aurelia","https://ap-avatar.wpscdn.com/avatar/10000253d8d9f28188e?_k=1776742907772140068",8,"Research & Report","CoCoT-EEG Contrastive Pretrained Multiscale Convolutional Transformer for EEG Decoding","Self-supervised foundation models are promising for non-invasive electroencephalogram (EEG) decoding, yet common masked reconstruction pretraining can be ineffective for EEG due to high noise amplitude and information concentrated in limited frequency bands. CoCoT-EEG introduces contrastive pretraining with multiscale temporal convolution input layers and Transformer encoders. The method matches or exceeds reconstruction-pretrained EEG models on heterogeneous electrode decoding benchmarks, and training from scratch also improves single-task baselines. Ablations validate contrastive learning feasibility and highlight key architectural choices for future EEG FM research.","CoCoT-EEG: Contrastive-Pretrained Multiscale Convolutional Transformer for EEG Decoding  \nGabriel Mahuas∗ Sigma Nova Paris, France  \n[gabriel.mahuas@sigmanova.ai](gabriel.mahuas@sigmanova.ai)  \nVictoria Shevchenko∗ Sigma Nova Paris, France  \n[victoria.shevchenko@sigmanova.ai](victoria.shevchenko@sigmanova.ai)  \nUgo Tanielian∗  \nSigma Nova  \nParis, France [ugo.tanielian@sigmanova.ai](ugo.tanielian@sigmanova.ai)  \nYassir Bendou  \nSigma Nova  \nParis, France [yassir.bendou@sigmanova.ai](yassir.bendou@sigmanova.ai)  \narXiv :2607 .09543v 1 [ cs .LG] 10 Jul 2026  \nRichard Gao  \nGoethe University Frankfurt  \nFrankfurt, Germany  \n[r.dg.gao@gmail.com](r.dg.gao@gmail.com)  \nAbstract  \nSelf-supervised pretrained foundation models (FM) have shown early promise for non-invasive electroencephalogram (EEG) decoding applications. Many recent large-scale models converged on the approach of tokenizing raw EEG followed by masked reconstruction pretraining. However, this recipe has been shown to be suboptimal for data—like EEG—with high noise amplitude and information confined to limited dimensions such as narrow frequency bands. Building on this insight, we develop a novel contrastive-pretrained EEG model with multiscale temporal convolution input layers and Transformer encoder blocks (CoCoT) . CoCoT matches or beats state-of-the-art reconstruction-pretrained EEG models on extensive benchmark decoding tasks with heterogeneous electrode configurations. Furthermore, CoCoT trained from scratch outperforms previous single-task decoding models and even rivals pretrained models, showcasing the architecture’s flexibility and data efficiency. Through systematic ablations, including model architecture and pretraining objective, we demonstrate the viability of contrastive learning for building EEG FMs while suggesting key architectural design considerations, prompting further investigations in alternative large-scale pretraining strategies.  \n1 Introduction  \nElectroencephalography (EEG) is pervasive in both research and clinical practice thanks to its moderate acquisition costs, non-invasive nature, and high temporal resolution. Decoding from EEG data remains an important challenge which, when addressed, enables myriad applications such as brain-computer interface (BCI) and clinical diagnosis. Classical deep learning-based solutions use deep neural networks in a supervised setting while extracting information from the time series, e.g., using temporal convolutions as the inductive bias of choice [1, 2, 3, 4] . While such solutions are generally effective, they are single-task models and often do not generalize well across settings [5] .  \n∗ Equal contribution.  \nPreprint.  \nRecently, large-scale self-supervised pretraining has proven successful at boosting cross-domain generalizability in language and computer vision. By leveraging vast amounts of unlabeled data, these models learn low-dimensional representations of their inputs that serve as generically useful features for downstream applications, often requiring little to no finetuning. Because of their general-purpose nature, such models have come to be known as foundation models (FMs) [6] .  \nThis pretrained FM strategy has also been adopted for EEG decoding: recent models like LaBraM [7], CBraMod [8], CSBrain [5], LUNA [9], and REVE [10] are pretrained on tens of thousands of hours of EEG data and perform well on a diverse set of downstream decoding tasks. Inspired by vision and language models, these solutions have converged on a particular recipe of early patch tokenization of the input data with reconstruction-based pretraining objectives (e.g., masked autoencoding, MAE [11]) . By further adapting this recipe to EEG data, such as discretizing into a codebook (via VQ-VAE [7]), dedicated spatial/temporal attention [8, 5], and more sophisticated electrode position encoding strategies [10, 5] combined with increasingly massive pretraining datasets [10], pretrained EEG FMshave emerged as potential alte","cbCaigVpU8xkO9zF","https://ap.wps.com/l/cbCaigVpU8xkO9zF","pdf",2260009,1,18,"English","en",105,"# Abstract\n# Introduction\n## Motivation and limitations of reconstruction pretraining\n## Contrastive learning for robust EEG representations\n# Main contributions","[{\"question\":\"What problem does the document address in EEG foundation model pretraining?\",\"answer\":\"The document targets limitations of reconstruction-based pretraining for EEG, especially in high-noise settings where information is concentrated in narrow frequency bands and early tokenization/reconstruction can be suboptimal.\"},{\"question\":\"What is CoCoT-EEG and how does it differ from prior reconstruction-pretrained approaches?\",\"answer\":\"CoCoT-EEG uses a convolution-first design with multiscale 1D convolutions followed by Transformer encoder blocks, and it applies contrastive pretraining objectives rather than masked reconstruction.\"},{\"question\":\"How does CoCoT-EEG perform compared with existing single-task and pretrained EEG models?\",\"answer\":\"The model trained from scratch outperforms previous best single-task approaches and can rival pretrained models. With contrastive pretraining, it further improves results across diverse downstream tasks, outperforming recent reconstruction-pretrained state of the art on the reported benchmarks.\"}]",1784180125,45,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"cocot-eeg-contrastive-pretrained-multiscale-convolutional-transformer-for-eeg-decoding","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/cocot-eeg-contrastive-pretrained-multiscale-convolutional-transformer-for-eeg-decoding/82396/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the document address in EEG foundation model pretraining?","Question",{"text":75,"@type":76},"The document targets limitations of reconstruction-based pretraining for EEG, especially in high-noise settings where information is concentrated in narrow frequency bands and early tokenization/reconstruction can be suboptimal.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What is CoCoT-EEG and how does it differ from prior reconstruction-pretrained approaches?",{"text":80,"@type":76},"CoCoT-EEG uses a convolution-first design with multiscale 1D convolutions followed by Transformer encoder blocks, and it applies contrastive pretraining objectives rather than masked reconstruction.",{"name":82,"@type":73,"acceptedAnswer":83},"How does CoCoT-EEG perform compared with existing single-task and pretrained EEG models?",{"text":84,"@type":76},"The model trained from scratch outperforms previous best single-task approaches and can rival pretrained models. With contrastive pretraining, it further improves results across diverse downstream tasks, outperforming recent reconstruction-pretrained state of the art on the reported benchmarks.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]