[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-124861-en":3,"doc-seo-124861-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},124861,549758252649,"Ivy","https://ap-avatar.wpscdn.com/avatar/8000253669c5317157?_k=1778319167496531819",8,"Research & Report","Enhanced EEG-Based Mental State Classification - A novel approach to eliminate data leakage and improve training optimization for Machine Learning","This paper investigates mental state level classification using electroencephalographic (EEG) signals and machine learning. It introduces an optimized training methodology that uses a validation set and a refined standardization procedure to address data leakage issues reported in earlier studies. The work also provides updated benchmark results across multiple models, including random forest and deep neural networks, enabling more reliable comparison and improved training optimization.","Enhanced EEG-Based Mental State Classification : A novel approach to eliminate data leakage and improve training optimization for Machine Learning  \nMaxime Girard, Rémi Nahon, Enzo Tartaglione and Van-Tam Nguyen  \nLTCI, Télécom Paris, Institut Polytechnique de Paris  \n[maxime.girard@telecom-paris.fr](maxime.girard@telecom-paris.fr)  \nAbstract—In this paper, we explore prior research and introduce a new methodology for classifying mental state levels based on EEG signals utilizing machine learning (ML). Our method proposes an optimized training method by introducing a validation set and a refined standardization process to rectify data leakage shortcomings observed in preceding studies. Furthermore, we establish novel benchmark figures for various models, including random forest and deep neural networks.  \nKeywords—Attention state classification, EEG, DNN, standardization  \nI. INTRODUCTION  \nAttention is the cognitive process involving the selection of environmental information for conscious processing [1] . Electroencephalographic (EEG) signals provide insights into the level of attention [2] . The surge in algorithms and machine learning models for attention level determination is driven by the diverse applications of attention classification. These applications span from educational tools incorporating workload monitoring to medical interventions for conditions like depression, schizophrenia, and anxiety disorders [3] . Recent studies employ signal analysis methods to leverage the frequency properties of brain signals [4][5] . Various models, such as LSTM [6], deep neural networks [4], or SVM [5][7], are trained to categorize frequencies derived from raw signals.  \nIn this paper, we draw upon the findings of previous research, specifically [4] and [5], to present a new approach to split a public dataset containing raw EEG signals for enhanced mental state classification. We delve into the standardization method outlined in [4] and propose an alternative technique to standardize the feature vector, addressing shortcomings inherent in the prior method. Ultimately, we define updated benchmark accuracy metricson the dataset employing various models, encompassing random forest and deep neural networks.  \nII. DATASET AND METHOD  \nA. Raw signals dataset used  \nWe utilized a publicly accessible dataset proposed in [5] . This dataset comprises raw signals captured from seven electrodes of a modified EEG Epoc headset placed on the brain cavity, identified as F3, F4, Fz, C3, C4, Cz, and Pz in the 10-20 electrode system [8] . The sampling frequency for the records is 128Hz. The dataset was formed from a population of five individuals. For each participant, except one, seven sessions have been recorded, each lasting at least 40 minutes. For the remaining participant, only six sessions were recorded. Each recording session consisted of three phases: in the initial ten minutes, participants were asked to control a train on a simulator (focused state); in the next ten minutes, participants were instructed to watch the screen  \nwithout controlling anything, with the instruction not to close their eyes (unfocused state); and after that, participants were asked to relax and close their eyes (drowsed state) . The first two records for each participant were used for habituation and are not suitable for classification [4] . It is important to note that the dataset's size is limited, which may pose a challenge for generalization [9] .  \nB. Previous works  \nPrior studies proposed the categorization of the dataset into three classes (focused, unfocused, drowsed) using diverse techniques [4][5] . Within these studies, three main paradigms for training a classifier can be identified:  \n1. Common-subject paradigm. The dataset is split into two parts, independently of the subject. 80% of the data is used for training, while the remaining 20% is reserved for evaluation.  \n2. Subject-specific paradigm. The classifier is trained on a single subject. For th","cbCaiaVipUiIPhVr","https://ap.wps.com/l/cbCaiaVipUiIPhVr","pdf",1025201,1,5,"English","en",105,"# Introduction\n# Dataset and Method\n## Raw signals dataset used\n## Previous works\n## Splitting method","[{\"question\":\"Why is data leakage a concern in prior EEG mental state classification studies?\",\"answer\":\"Earlier approaches split data into train and test without the proposed safeguards, which can allow unintended information overlap. The paper addresses this by introducing a validation set and a refined standardization process.\"},{\"question\":\"What EEG dataset and recording setup are used for the experiments?\",\"answer\":\"The study uses a public dataset recorded from seven electrodes (F3, F4, Fz, C3, C4, Cz, Pz) with a sampling frequency of 128 Hz from five individuals across multiple sessions and phases.\"},{\"question\":\"How does the proposed dataset splitting method differ from common paradigms?\",\"answer\":\"The method focuses on the leave-one-out paradigm but adds a validation set and standardizes features using an alternative technique, splitting data into train, validation, and test rather than only train/test.\"}]","Enhanced EEG-Based Mental State Classification - A novel approach to eliminate data leakage and improve training optimization for Machine Learning | PDF",1785895085,13,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"enhanced-eeg-based-mental-state-classification-a-novel-approach-to-eliminate-data-leakage-and-improve-training-optimization-for-machine-learning","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/enhanced-eeg-based-mental-state-classification-a-novel-approach-to-eliminate-data-leakage-and-improve-training-optimization-for-machine-learning/124861/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-05",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why is data leakage a concern in prior EEG mental state classification studies?","Question",{"text":75,"@type":76},"Earlier approaches split data into train and test without the proposed safeguards, which can allow unintended information overlap. The paper addresses this by introducing a validation set and a refined standardization process.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What EEG dataset and recording setup are used for the experiments?",{"text":80,"@type":76},"The study uses a public dataset recorded from seven electrodes (F3, F4, Fz, C3, C4, Cz, Pz) with a sampling frequency of 128 Hz from five individuals across multiple sessions and phases.",{"name":82,"@type":73,"acceptedAnswer":83},"How does the proposed dataset splitting method differ from common paradigms?",{"text":84,"@type":76},"The method focuses on the leave-one-out paradigm but adds a validation set and standardizes features using an alternative technique, splitting data into train, validation, and test rather than only train/test.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,109,114,119,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":21,"doc_module":4,"doc_module_name":46,"category_name":106,"show_sort_weight":107,"slug":108},"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":46,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":21,"slug":137},19,"General","general"]