[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-122472-en":3,"doc-seo-122472-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},122472,1374391974585,"Genevieve","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","Enhancing Soil Liquefaction Prediction - Overcoming Data Challenges in SPT-Based Machine Learning with Imputation","Earthquakes not only cause direct damage but also reduce soil-bearing capacity during liquefaction, worsening failures of buildings and infrastructure. Liquefaction depends on many interacting parameters, making evaluation difficult, especially when datasets contain incomplete liquefaction records. This study evaluates machine learning model capability for predicting liquefaction by applying missing value imputation. Seismicity, soil properties, and soil condition parameters from standard penetration test (SPT) data support Random Forest, k-NN, and XGBoost training with feature selection and parameter optimization, while normalization and outlier treatment improve prediction reliability. Confusion-matrix metrics and AUC show Random Forest achieves the highest overall accuracy (OA=90.71%), indicating strong classification performance. Overall, imputation combined with preprocessing substantially enhances data-driven geotechnical earthquake modeling.","Journal of the Civil Engineering Forum, January 2026, 12(1):23-39  \nDOI 10.22146/jcef.21347  \nAvailable Online at [https://jurnal.ugm.ac.id/v3/jcef/issue/archive](https://jurnal.ugm.ac.id/v3/jcef/issue/archive)  \nEnhancing Soil Liquefaction Prediction: Overcoming Data Challenges in SPT-Based Machine Learning with Imputation  \nTechnique  \nFandi Fadliansyah1 , Fikri Faris1,4* , Wahyu Wilopo2,4 , Ardiansyah3  \n1 Department of Civil and Environmental Engineering, Universitas Gadjah Mada, Yogyakarta, INDONESIA  \n2 Department of Geological Engineering, Universitas Gadjah Mada, Yogyakarta, INDONESIA  \n3 Department of Computer Science, Faculty of Mathematics and Natural Sciences, Universitas Lampung, Lampung, INDONESIA  \n4 Center for Disaster Mitigation and Technological Innovation (GAMA-InaTEK), Universitas Gadjah Mada, Yogyakarta, INDONESIA  \n*[Corresponding author: fikri.faris@ugm.ac.id](Corresponding author: fikri.faris@ugm.ac.id)  \nSUBMITTED 05 May 2025 REVISED 20 July 2025 ACCEPTED 23 July 2025  \nABSTRACT In addition to the adverse effects of earthquakes, the loss of soil-bearing capacity during liquefaction can exacerbate damage to buildings. Liquefaction phenomena involve many parameters, making it more complex to evaluate. Machine learning has been studied to deal with liquefaction complexity in recent decades. However, incomplete liquefaction data can result in missing information, complicating model development across various datasets. Therefore, this study aims to assess the capability of machine learning models to predict liquefaction by implementing the missing value imputation technique. Seismicity, soil properties, and soil condition parameters were utilized to develop models. Random Forest (RF), k-Nearest Neighbor (k-NN), and eXtreme Gradient Boosting (XGBoost) were trained by applying feature selection and parameter optimization based on standard penetration test (SPT) data. The confusion matrix was used to assess the performance of the model based on the performance matrix of Overall Accuracy (OA), Precision (Prec), Recall (Rec), F1-Score (F1), and Area Under the Curve (AUC) . In addition, the preprocessing stage included data normalization and outlier treatment to enhance the reliability of model predictions, ensuring consistent learning behavior across different variable scales. The results show that the RF achieved the highest performance (OA = 90.71%), which is comparable to findings from other previous studies. The AUC results indicate that the models deliver excellent classification performance. These findings suggest that the integration of imputation and preprocessing techniques can significantly improve data-driven approaches in geotechnical earthquake engineering. In conclusion, the missing imputation is quite effective in the predictive model. Finally, this study offers a new perspective on developing machine learning models using a more user-friendly software and applying imputation techniques to handle missing data.  \nKEYWORDS Machine learning; Missing value imputation; Soil liquefaction; Earthquake; Standard penetration test.  \n© The Author(s) 2026 . This article is distributed under a Creative Commons Attribution-ShareAlike 4 .0 International license.  \n1 INTRODUCTION  \n1.1 Seismic-Induced Liquefaction and Advancesin Identification Methods  \nThe movement of tectonic plates can lead to vibrations known as earthquakes. The intensity of an earthquake can lead to devastating effects. In addition, earthquakes, earthquakes can trigger other natural disasters, such as soil liquefaction. Soil liquefaction can worsen the damage to building infrastructure. This phenomenon caused extensive damage, as happened in the 1964 Niigata earthquake in Japan, the 1964 Alaska earthquake, the 1999 Chi-Chi and Kocaeli earthquakes, and the 2018 Palu earthquake. Soil liquefaction can cause the building’s foundation to crack and collapse, resulting in fatalities. As a result of liquefaction, formerly solid soil transforms into ","cbCaiuQJ5vVQ219p","https://ap.wps.com/l/cbCaiuQJ5vVQ219p","pdf",2653782,1,17,"English","en",105,"# Introduction\n## Seismic-Induced Liquefaction and Advances in Identification Methods\n## Overview of Machine Learning Models in Liquefaction Assessment","[{\"question\":\"Why is soil liquefaction prediction challenging for machine learning models?\",\"answer\":\"Liquefaction involves many parameters and often lacks complete data. Missing liquefaction records create missing information that complicates model development across different datasets.\"},{\"question\":\"What missing-data strategy does the study use?\",\"answer\":\"The study applies missing value imputation as part of the preprocessing pipeline. This enables machine learning models to learn effectively despite incomplete liquefaction data.\"},{\"question\":\"Which model achieved the best prediction performance and how was it evaluated?\",\"answer\":\"Random Forest achieved the highest performance with OA=90.71%. Model quality was assessed using confusion-matrix metrics such as Precision, Recall, F1-Score, and AUC, alongside preprocessing such as normalization and outlier treatment.\"}]","Enhancing Soil Liquefaction Prediction - Overcoming Data Challenges in SPT-Based Machine Learning with Imputation | PDF",1785810835,43,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"enhancing-soil-liquefaction-prediction-overcoming-data-challenges-in-spt-based-machine-learning-with-imputation","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/enhancing-soil-liquefaction-prediction-overcoming-data-challenges-in-spt-based-machine-learning-with-imputation/122472/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why is soil liquefaction prediction challenging for machine learning models?","Question",{"text":75,"@type":76},"Liquefaction involves many parameters and often lacks complete data. Missing liquefaction records create missing information that complicates model development across different datasets.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What missing-data strategy does the study use?",{"text":80,"@type":76},"The study applies missing value imputation as part of the preprocessing pipeline. This enables machine learning models to learn effectively despite incomplete liquefaction data.",{"name":82,"@type":73,"acceptedAnswer":83},"Which model achieved the best prediction performance and how was it evaluated?",{"text":84,"@type":76},"Random Forest achieved the highest performance with OA=90.71%. Model quality was assessed using confusion-matrix metrics such as Precision, Recall, F1-Score, and AUC, alongside preprocessing such as normalization and outlier treatment.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]