[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-127659-en":3,"doc-seo-127659-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},127659,962084925636,"Olivia Brown","https://ap-avatar.wpscdn.com/davatar_994ba38a5ba835b3df7d355c54d3ed8d",8,"Research & Report","Exploring Statistical and Machine Learning-Based Missing Data Imputation Methods to Improve Crash Frequency Prediction Models for Highway-Rail Grade Crossings - Paper","Highway-rail grade crossings (HRGCs) are high-risk transportation safety locations where crashes can result in severe injuries and fatalities. In the United States, large numbers of crashes motivate organizations to estimate crash frequency prediction models to support safety planning and risk mitigation. Model quality depends on HRGC inventory data, yet many models exclude crossings with missing inventory details, harming prediction precision. This study analyzes a filtered sample of 2000 HRGCs, imputes missing values using statistical and machine learning methods, and evaluates model fit via AIC and BIC.","Mid-America Transportation Center: Final Reports and Technical Briefs  \nMid-America Transportation Center  \n11-14-2023  \nExploring Statistical and Machine Learning-Based Missing Data Imputation Methods to Improve Crash Frequency Prediction Models for Highway-Rail Grade Crossings  \nMuhammad Umer Farooq Aemal Khattak  \nFollow this and additional works at: [https://digitalcommons.unl.edu/matcreports](https://digitalcommons.unl.edu/matcreports)  \n Part of the Civil Engineering Commons, and the Transportation Engineering Commons  \nThis Article is brought to you for free and open access by the Mid-America Transportation Center at DigitalCommons@University of Nebraska-Lincoln. It has been accepted for inclusion in Mid-America Transportation Center: Final Reports and Technical Briefs by an authorized administrator of DigitalCommons@University of Nebraska-Lincoln.  \nPhoenix, AZ  \n\n| PAPER TITLE | Exploring Statistical and Machine Learning-Based Missing Data Imputation Methods to Improve Crash Frequency Prediction Models for Highway-Rail\u003Cbr>Grade Crossings |  |  |\n| --- | --- | --- | --- |\n| TRACK |  |  |  |\n| AUTHOR (Capitalize Family Name) | POSITION | ORGANIZATION | COUNTRY |\n| Muhammad Umer FAROOQ | Post Doctoral Research Associate, MidAmerica Transportation Center (MATC) | University of Nebraska-Lincoln | USA |\n| CO-AUTHOR(S)(Capitalize Family Name | POSITION | ORGANIZATION | COUNTRY |\n| Aemal KHATTAK | Professor and Director, Mid-America Transportation Center (MATC) | University of Nebraska-Lincoln | USA |\n| E-MAIL\u003Cbr>(for correspondence) | [mfarooq2@unl.edu](mfarooq2@unl.edu) |  |  |\n| KEYWORDS:\u003Cbr>Missing data imputation, Crash frequency prediction, Highway-rail grade crossings, Statistical methods, Machine learning-based methods |  |  |  |\n\nABSTRACT:  \nHighway-rail grade crossings (HRGCs) are critical spatial locations of transportation safety because crashes at HRGCs are often catastrophic, potentially causing several injuries and fatalities. Every year in the United States, a significant number of crashes occur at these crossings, prompting local and state organizations to engage in safety analysis and estimate crash frequency prediction models for resource allocation. These models provide valuable insights into safety and risk mitigation strategies for HRGCs. Furthermore, the estimation of these models is based on inventory details of HRGCs, and their quality is crucial for reliable crash predictions. However, many of these models exclude crossings with missing inventory details, which can adversely affect the precision of these models. In this study, a random sample of inventory details of 2000 HRGCs was taken from the Federal Railroad Administration’s HRGCs inventory database. Data filters were applied to retain only those crossings in the data that were at-grade, public and operational (N=1096) . Missing values were imputed using various statistical and machine learning methods, including Mean, Median and Mode (MMM) imputation, Last Observation Carried Forward (LOCF) imputation, K-Nearest Neighbors (KNN) imputation, Expectation-Maximization (EM) imputation, Support Vector Machine (SVM) imputation, and Random Forest (RF) imputation. The results indicated that the crash frequency models based on machine learning imputation methods yielded better-fitted models (lower AIC and BIC values) . The findings underscore the importance of obtaining complete inventory data through machine learning imputation methods when developing crash frequency models for HRGCs. This approach can substantially enhance the precision of these models, improving their predictive capabilities, and ultimately saving valuable human lives.  \nExploring Statistical and Machine Learning-Based Missing Data Imputation Methods to Improve Crash Frequency Prediction Models for  \nHighway-Rail Grade Crossings  \nDr. Muhammad Umer Farooq and Dr. Aemal Khattak 1  \n1Mid-America Transportation Center, University of Nebraska Lincoln, USA  \nEmail for correspondence, [e.g.]","cbCaihzYpw19Hgk0","https://ap.wps.com/l/cbCaihzYpw19Hgk0","pdf",585085,1,14,"English","en",105,"# Abstract\n# Introduction\n## Missing data imputation in predictive modeling\n## Impact of missing values on model accuracy\n## Overview of statistical and machine learning approaches","[{\"question\":\"Why is missing data imputation important for crash frequency prediction models at HRGCs?\",\"answer\":\"Missing values create gaps, reducing sample size and introducing bias, noise, and distorted statistical properties. Imputation enables using available information rather than discarding incomplete records, improving predictive reliability.\"},{\"question\":\"Which imputations were evaluated in the study?\",\"answer\":\"The study evaluated MMM (Mean, Median, Mode), LOCF, KNN, Expectation-Maximization (EM), SVM-based imputation, and Random Forest (RF) imputation.\"},{\"question\":\"How were the models compared to determine which imputations performed best?\",\"answer\":\"The fitted crash frequency models were evaluated using AIC and BIC values, with lower values indicating better model fit.\"}]","Exploring Statistical and Machine Learning-Based Missing Data Imputation Methods to Improve Crash Frequency Prediction Models for Highway-Rail Grade Crossings - Paper | PDF",1785940567,35,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"exploring-statistical-and-machine-learning-based-missing-data-imputation-methods-to-improve-crash-frequency-prediction-models-for-highway-rail-grade-crossings-paper","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/exploring-statistical-and-machine-learning-based-missing-data-imputation-methods-to-improve-crash-frequency-prediction-models-for-highway-rail-grade-crossings-paper/127659/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-23","2026-08-05",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"Why is missing data imputation important for crash frequency prediction models at HRGCs?","Question",{"text":76,"@type":77},"Missing values create gaps, reducing sample size and introducing bias, noise, and distorted statistical properties. Imputation enables using available information rather than discarding incomplete records, improving predictive reliability.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"Which imputations were evaluated in the study?",{"text":81,"@type":77},"The study evaluated MMM (Mean, Median, Mode), LOCF, KNN, Expectation-Maximization (EM), SVM-based imputation, and Random Forest (RF) imputation.",{"name":83,"@type":74,"acceptedAnswer":84},"How were the models compared to determine which imputations performed best?",{"text":85,"@type":77},"The fitted crash frequency models were evaluated using AIC and BIC values, with lower values indicating better model fit.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":46,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":46,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]