[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-119038-en":3,"doc-seo-119038-105":30,"detail-sidebar-cat-0-en-105":83},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},119038,137441390410,"Hazel","https://ap-avatar.wpscdn.com/avatar/2000252f4ab5702993?_k=1776741390130283984",6,"Technology","Scope and Arbitration in Machine Learning Clinical EEG Classification","Clinical EEG interpretation depends on classifying an entire recording or visit as normal versus abnormal. Machine learning pipelines often split recordings into short windows for feasibility, then propagate the parent session label to every window, which can introduce misleading supervision. This study tests two fixes: longer windows and a second-stage arbitration model that reconciles window-level predictions within each recording. On the Temple University Hospital Abnormal EEG Corpus, average accuracy improves from 89.8% to 93.3%, exceeding prior performance expectations.","Scope and Arbitration in Machine Learning Clinical EEG  \nClassiﬁcation  \narXiv :2303 .06386v1 [ cs .LG] 11 Mar 2023  \nYixuan Zhu [yixuan2.zhu@live.uwe.ac.uk](yixuan2.zhu@live.uwe.ac.uk)  \nUniversity of the West of England, UK  \nLuke J. W. Canham luke .canham@nbt.nhs .uk  \nNorth Bristol NHS Trust, UK  \nDavid Western [david.western@uwe.ac.uk](david.western@uwe.ac.uk)  \nUniversity of the West of England, UK  \nAbstract  \nA key task in clinical EEG interpretation is to classify a recording or session as normal or abnormal. In machine learning approaches to this task, recordings are typically divided into shorter windows for practical reasons, and these windows inherit the label of their parent recording. We hypothesised that window labels derived in this manner can be misleading  \n– for example, windows without evident abnormalities can be labelled `abnormal' – disrupting the learning process and degrading performance. We explored two separable approaches to mitigate this problem: increasing the window length and introducing a second-stage model to arbitrate between the window-speciﬁc predictions within a recording. Evaluating these methods on the Temple University Hospital Abnormal EEG Corpus, we signiﬁcantly improved state-of-the-art average accuracy from 89.8 percent to 93.3 percent. This result deﬁes previous estimates of the upper limit for performance on this dataset and represents a major step towards clinical translation of machine learning approaches to this problem.  \nData and Code Availability Our study includes electroencephalography (EEG) datasets collected from [https://isip.piconepress.com/projects/](https://isip.piconepress.com/projects/)[ ](https://isip.piconepress.com/projects/)tuh_eeg/. Our code is shared on [https://github](https://github). com/zhuyixuan1997/EEGScopeAndArbitration.  \n1. Introduction  \n1.1. Background  \nElectroencephalography (EEG) recordings are used for the diagnosis and monitoring of a wide range of neurological conditions. Classiﬁcation of EEG  \nrecordings as normal or abnormal is an essential task in their clinical interpretation. Substantial research has been conducted on the application of machine learning to this task (Schirrmeister et al. , 2017 ; Amin et al. , 2019 ; Banville et al. , 2021 , 2022 ; Gemein et al. , 2020 ; Muhammad et al. , 2020 ; Wagh and Varatharajah, 2020 ; Alhussein et al. , 2019 ; Roy et al. , 2019) .  \nRecent work in this ﬁeld largely makes use of the Temple University Hospital Abnormal EEG Corpus (TUAB) (López et al. , 2017) for training and evaluation. TUAB is a labelled subset of the Temple University Hospital EEG Corpus (TUEG) (Obeid and Picone, 2016) .  \nSince the presentation of the Deep4 convolutional neural network in 2017 (Schirrmeister et al. , 2017) there have been only modest improvements in the accuracy of machine learning approaches to this problem, as measured on TUAB: from 85.4 percent (Deep4) up to 89.8 percent (Muhammad et al. , 2020)  \n– see Table 1 for further detail. Gemein et al. (2020) proposed that there may be an upper limit of around 90 percent accuracy in this task, based on known values of inter-rater agreement between human experts in conventional clinical practice.  \nA notable but little-discussed diﬀerence between conventional clinical practice and virtually all deep learning approaches is that in clinical practice, the label of normal/abnormal is applied to a full EEG session (i.e. a single clinical visit) . In clinical practice, experts judge whether the patient exhibits abnormal brain activity based on all the recordings in the session, eﬀectively resulting in a single label for that session. In most recent machine learning approaches, a typical full recording cannot be directly input into the model due to computational constraints – a large input vector length necessitates a large number of  \n© Y. Zhu, L.J.W. Canham & D. Western.  \nScope and Arbitration in EEG Classification  \nTable 1: Summary of state-of-the-art performance metric","cbCaipYHOkofCdzO","https://ap.wps.com/l/cbCaipYHOkofCdzO","pdf",2485735,1,10,"English","en",105,"# Introduction\n## Background\n## Proposal\n### Overview\n# Data and Code Availability","[{\"question\":\"How much does the proposed method improve performance on the TUAB dataset?\",\"answer\":\"Average accuracy improves from 89.8% to 93.3%, outperforming the previous reported upper-limit estimate near 90% for this dataset.\"}]","Scope and Arbitration in Machine Learning Clinical EEG Classification | PDF",1785722037,25,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":78,"head_meta":80,"extra_data":82,"updated_unix":28},"scope-and-arbitration-in-machine-learning-clinical-eeg-classification","",{"@graph":36,"@context":77},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/technology/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/scope-and-arbitration-in-machine-learning-clinical-eeg-classification/119038/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71],{"name":72,"@type":73,"acceptedAnswer":74},"How much does the proposed method improve performance on the TUAB dataset?","Question",{"text":75,"@type":76},"Average accuracy improves from 89.8% to 93.3%, outperforming the previous reported upper-limit estimate near 90% for this dataset.","Answer","https://schema.org",{"og:url":52,"og:type":79,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":81,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":84},[85,89,93,97,102,105,110,115,120,123,126],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":86,"show_sort_weight":87,"slug":88},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":90,"show_sort_weight":91,"slug":92},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Exam",70,"exam",{"id":98,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},5,"Comic",60,"comic",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":103,"slug":104},50,"technology",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},7,"Healthcare",40,"healthcare",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},8,"Research & Report",30,"research-report",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},9,"Religion & Spirituality",20,"religion-spirituality",{"id":118,"doc_module":4,"doc_module_name":46,"category_name":121,"show_sort_weight":118,"slug":122},"World Cup","world-cup",{"id":21,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":21,"slug":125},"Lifestyle","lifestyle",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":98,"slug":129},19,"General","general"]