[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-120118-en":3,"doc-seo-120118-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},120118,8796095462418,"Noah","https://ap-avatar.wpscdn.com/avatar/80000253c1241d02b47?x-image-process=image/resize,m_fixed,w_180,h_180&k=1778826106357471780",8,"Research & Report","Machine-learning-based particle identification with missing data - paper","This work presents a new Particle Identification (PID) method for the ALICE experiment at the Large Hadron Collider, targeting reliable identification of collision products. Traditional PID relies on hand-crafted selections compared with simulations, while modern machine learning classifiers require complete detector signals. ALICE subdetectors may record only subsets of measurements due to inefficiencies, acceptance limits, malfunctions, or kinematic mismatches, creating missing values. The proposed architecture enables training on both complete and incomplete examples, improving PID purity and efficiency across investigated particle species.","arXiv :2401 .01905v2 [physics .ins-det] 22 Jul 2024  \nMachine-learning-based particle identification with missing data  \nMilosz Kasaka , Kamil Dejaa,b , Maja Karwowskaa,c , Monika Jakubowskaa , Lukasz Graczykowskia , Malgorzata Janikaa Warsaw University of Technology, pl. Politechniki 1, 00-661 , Warsaw,  \nPoland.  \nb IDEAS NCBR, Chmielna 69, 00-801, Warsaw, Poland.  \nc CERN – European Organization for Nuclear Research, Espl. des Particules 1, 1211 Geneva, Switzerland.  \nAbstract  \nIn this work, we introduce a novel method for Particle Identification (PID) within the scope of the ALICE experiment at the Large Hadron Collider at CERN. Identifying products of ultrarelativisitc collisions delivered by the LHC is oneof the crucial objectives of ALICE. Typically employed PID methods rely on hand-crafted selections, which compare experimental data to theoretical simulations. To improve the performance of the baseline methods, novel approaches use machine learning models that learn the proper assignment in a classification task. However, because of the various detection techniques used by different subdetectors, as well as the limited detector efficiency and acceptance, produced particles do not always yield signals in all of the ALICE components. This results in data with missing values. Out of the box machine learning solutions cannot be trained with such examples without either modifying the training dataset or re-designing the model architecture. In this work, we propose the new method for PID that addresses these issues and can be trained with all of the available data examples, including incomplete ones. Our approach improves the PID purity and efficiency of the selected sample for all investigated particle species.  \nKeywords: particle identification, machine learning, missing data  \n1  \nFig. 1: Components of the ALICE detector in its Run 2 configuration [4] .  \n1 Introduction  \nALICE (A Large Ion Collider Experiment) [1] is one of the four major detectors located at the Large Hadron Collider at CERN [2] . The main goal of ALICE is to study the properties of quark–gluon plasma (QGP), a hot and dense state of matter, and the strong force that binds quarks together inside hadrons [3] . The key requirement for detailed studies of QGP that distinguishes ALICE from the other Large Hadron Collider (LHC) experiments is its capability for very precise particle identification (PID)– i.e. the ability to discriminate between different particle species produced during the collision. This allows for selecting a subset of particles required for specific analysis.  \nThe ALICE experiment is composed of several sub-detectors, some of which measure particle properties that can be used for identification. Figure 1 presents a scheme of the detector in Run 2, the previous LHC data-taking periods.  \nThe three detectors particularly useful for PID are: TPC, TOF, and TRD. The Time Projection Chamber [5](TPC) is one of the most important ALICE detectors as it records 3D information of the trajectory of charged particles, as well as their specific energy loss due to ionization, which is essential for particle identification. The Time-of-Flight [6](TOF) detector measures particle travel times from the collision vertex to the detector, from which the particle velocity and mass are calculated. The Transition Radiation Detector [7] (TRD) records transition radiation, that is, the emission of photons by electrons traversing the boundaries of a radiator, which helps in distinguishing electrons from other charged particles. All the detectors mentioned above detect particles carrying a non-zero electric charge. Therefore, this article focuses on the identification of charged particles.  \nWith the signals recorded by the detectors described above, particles are chosen using a set of selection criteria. Traditionally, the particle identification is based on hand-crafted selection criteria, for instance, based on how much the detector  \n2  \nresponse deviates from","cbCaijMbcUBDfLW5","https://ap.wps.com/l/cbCaijMbcUBDfLW5","pdf",22308645,1,23,"English","en",105,"# Introduction\n## ALICE experiment and PID goal\n## PID-relevant subdetectors (TPC, TOF, TRD)\n## Limitations of hand-crafted and standard ML methods with missing data\n# Proposed method (overview)\n## Training with incomplete detector measurements\n## Expected impact on PID purity and efficiency","[{\"question\":\"Why does ALICE PID training data contain missing values?\",\"answer\":\"Subdetectors used for PID may not record signals for every produced particle due to limited efficiency/acceptance, detector issues, or particle properties that fall outside detector specifications.\"},{\"question\":\"What limits out-of-the-box machine learning solutions for PID with missing data?\",\"answer\":\"Standard models typically require complete input features, so they cannot be trained or applied directly when some detector measurements are absent.\"},{\"question\":\"How does the proposed approach address missing data in PID?\",\"answer\":\"It modifies the model architecture so training can use all available examples, including incomplete ones, without assuming how missing values should be imputed.\"}]","Machine-learning-based particle identification with missing data - paper | PDF",1785728302,58,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"machine-learning-based-particle-identification-with-missing-data-paper","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/machine-learning-based-particle-identification-with-missing-data-paper/120118/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why does ALICE PID training data contain missing values?","Question",{"text":75,"@type":76},"Subdetectors used for PID may not record signals for every produced particle due to limited efficiency/acceptance, detector issues, or particle properties that fall outside detector specifications.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What limits out-of-the-box machine learning solutions for PID with missing data?",{"text":80,"@type":76},"Standard models typically require complete input features, so they cannot be trained or applied directly when some detector measurements are absent.",{"name":82,"@type":73,"acceptedAnswer":83},"How does the proposed approach address missing data in PID?",{"text":84,"@type":76},"It modifies the model architecture so training can use all available examples, including incomplete ones, without assuming how missing values should be imputed.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]