[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-124500-en":3,"doc-seo-124500-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},124500,13056703020460,"Valentina","https://ap-avatar.wpscdn.com/avatar/be000253dac470eee5d?_k=1778207105932848923",8,"Research & Report","Prediction of Stark Broadening of Atomic Spectral Lines Using Machine Learning - Bachelor’s Thesis","Stark broadening and spectral-line shifts are essential parameters for diagnosing astrophysical and laboratory plasmas, but established semi-classical or semi-empirical theoretical methods are costly and limited in coverage. This thesis proposes a data-driven machine learning pipeline using large-scale STARK-B database data to predict Stark parameters rapidly and accurately. A deep feature engineering approach integrates physical prior knowledge by parsing spectroscopic text into meaningful numerical features. Trained on 130,000+ processed points, the best model achieves R²=0.952 for Stark widths and 5.71% median relative error.","YILIN WANG  \nPrediction of Stark Broadening of  \nAtomic Spectral Lines Using Machine Learning  \nDEGREE PROGRAMME IN DATA ENGINEERING  \n2025  \nABSTRACT  \nWang , Yilin: Prediction of Stark Broadening of Atomic Spectral Lines Using  \nMachine Learning Bachelor’s thesis Data Engineering November 2025 Number of pages: 68  \nStark broadening and shift of atomic spectral lines are fundamental parameters for the diagnosis of astrophysical and laboratory plasmas. Traditional theoretical calculation methods, such as Semi-Classical Perturbation (SCP) or Modified Semi-Empirical (MSE) approaches, are computationally expensive and often limited in coverage. This paper presents a data-driven approach based on machine learning to achieve rapid and high-precision prediction of Stark parameters using large-scale data from the STARK-B database. The core innovation of this study lies in the development of a deep feature engineering pipeline integrated with physical prior knowledge. Addressing the complex textual representations in raw spectroscopic data, I developed specialized parsing algorithms to extract physically meaningful numerical features—such as orbital occupancy counts, multiplicity, L-quantum numbers, and parity—from electron configurations, spectral terms, and total angular momentum (J) .  \nOver 130,000 processed data points were used to evaluate the performance of various models, including Random Forest, Neural Network, and XGBoost. The final model demonstrated exceptional accuracy for Stark widths, achieving a Coefficient of Determination (􀜴2 ) of 0.952 and a Median Relative Error (MdRE) of just 5.71% on an independent test set. While Stark shift predictions proved more challenging due to near-zero values, the model successfully captured macroscopic trends (􀜴2 ≈ 0. 92) . This research confirms that machine learning, when combined with deep physical feature extraction, offers a scalable and reliable alternative for populating large-scale atomic spectral databases.  \nKeywords: Stark Effect , Machine Learning , Feature Engineering , XGBoost , Atomic Spectroscopy , Plasma Diagnosis  \nPREFACE  \nThis thesis was undertaken as part of my graduation project at SAMK. The topic was kindly assigned by Toni Aaltonen, who provided continuous guidance throughout the research process. Numerous discussions and meetings with Toni Aaltonen significantly contributed to the development and refinement of the ideas presented in this thesis. Over the past two and a half years, Toni has offered much help in my studies, for which I am sincerely grateful.  \nI would also like to thank Vi, Hailin and Runying for their support, companionship, and encouragement during my time at SAMK. Studying at SAMK proved to be a genuinely fulfilling experience, one that fostered both academic learning and personal development.  \nLooking back, my study journey at SAMK happened to take place during the explosive rise of artificial intelligence. ChatGPT was released only a few months before I started my studies, and throughout these years I have witnessed how rapidly AI has entered, integrated into, and transformed our daily lives. This technological transformation also directly affected this thesis; despite having almost zero background in physics, I was able to embark on this thesis simply because I was interested in the topic. AI helped bridge the gapsin my theoretical knowledge and made it possible for me to explore a field that would otherwise have been out of reach. For this, I feel very lucky about this, and I am grateful to the predecessors in the AI field. I also hope to contribute my efforts in this magnificent process.  \nFinally, I hope this work will contribute, in some small way, to the understanding and prediction of Stark broadening phenomena, and demonstrate the potential of combining traditional physics with modern computational methods.  \nYilin Wang Nov 2025  \nCONTENTS  \n1 INTRODUCTION .............................................................................","cbCaipg7Rc4ONu2p","https://ap.wps.com/l/cbCaipg7Rc4ONu2p","pdf",3249868,1,67,"English","en",105,"# 1 INTRODUCTION\n# 2 THEORETICAL BACKGROUND OF STARK BROADENING\n## 2.1 The Quantum Mechanical Basis of the Atomic Model\n## 2.2 Radiative Transitions and Line Formation\n## 2.3 Perturbation by Electric Fields: The Stark Effect\n## 2.4 Statistical Broadening\n# 3 DATA PREPARATION\n## 3.1 Data Source\n## 3.2 Data Acquisition Pipeline\n## 3.3 Data Transformation\n## 3.4 Merge Data\n# 4 DATA CLEANING AND QUALITY ASSURANCE\n## 4.1 Dimensionality Reduction via Sparsity Filtering\n## 4.2 Resolution of Parsing Anomalies\n## 4.3 Missing Value Handling and Data Validation\n# 5 FEATURE ENGINEERING","[{\"question\":\"What problem does the thesis address?\",\"answer\":\"It addresses the need for accurate and fast prediction of Stark broadening and spectral-line shifts, which are key parameters for diagnosing plasma environments.\"},{\"question\":\"How is machine learning applied in the study?\",\"answer\":\"The work trains multiple models, including Random Forest, Neural Network, and XGBoost, using large-scale processed data from the STARK-B database.\"},{\"question\":\"What makes the feature engineering approach special?\",\"answer\":\"It uses deep feature engineering combined with physical prior knowledge, extracting physically meaningful numerical features from electron configurations, spectral terms, and quantum numbers.\"}]","Prediction of Stark Broadening of Atomic Spectral Lines Using Machine Learning - Bachelor’s Thesis | PDF",1785822791,169,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"prediction-of-stark-broadening-of-atomic-spectral-lines-using-machine-learning-bachelors-thesis","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/prediction-of-stark-broadening-of-atomic-spectral-lines-using-machine-learning-bachelors-thesis/124500/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the thesis address?","Question",{"text":75,"@type":76},"It addresses the need for accurate and fast prediction of Stark broadening and spectral-line shifts, which are key parameters for diagnosing plasma environments.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How is machine learning applied in the study?",{"text":80,"@type":76},"The work trains multiple models, including Random Forest, Neural Network, and XGBoost, using large-scale processed data from the STARK-B database.",{"name":82,"@type":73,"acceptedAnswer":83},"What makes the feature engineering approach special?",{"text":84,"@type":76},"It uses deep feature engineering combined with physical prior knowledge, extracting physically meaningful numerical features from electron configurations, spectral terms, and quantum numbers.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]