[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-126309-en":3,"doc-seo-126309-105":31,"detail-sidebar-cat-0-en-105":93},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":30},126309,2336475104957,"Seraphina","https://ap-avatar.wpscdn.com/avatar/22000c4c6bd8a5076e1?x-image-process=image/resize,m_fixed,w_180,h_180&k=1787554080175789136",8,"Research & Report","Identification of Potentially Misclassified Crash Narratives using Machine Learning (ML) and Deep Learning (DL)","This research investigates the efficacy of machine learning (ML) and deep learning (DL) methods in detecting misclassified intersection-related crashes within police-reported narratives. Using 2019 crash data from the Iowa Department of Transportation, the study implements and compares SVM, XGBoost, BERT sentence and word embeddings, and Albert-based models. Performance is validated against expert reviews of potentially misclassified narratives, ensuring rigorous classification accuracy. Results show hybrid strategies that combine narrative text with structured crash data reduce error rates by 54.2%, improving crash data quality for transportation safety management.","arXiv :2507 .03066v 1 [ cs .CL] 3 Jul 2025  \nIdentification of Potentially Misclassified  \nCrash Narratives  \nusing Machine Learning (ML) and Deep  \nLearning (DL)  \nSudesh Bhagat*1 , Ibne Farabi Shihab*2 , and Jonathan Wood3  \n1 Department of Civil, Construction and Environmental Engineering, , Iowa State University, 813 Bissell Road, Ames, IA 50011, USA. ,  \n[bhagat@iastate. edu](bhagat@iastate. edu)  \n2 Department of Computer Science, , Iowa State University, Ames, IA 50011, USA. , [ishihab@iastate. edu](ishihab@iastate. edu)  \n3 Department of Civil, Construction and Environmental Engineering, , Iowa State University, 813 Bissell Road, Ames, IA 50011, USA. ,  \n[jwood2@iastate. edu](jwood2@iastate. edu)  \nAbstract  \nThis research investigates the efficacy of machine learning (ML) and deep learning (DL) methods in detecting misclassified intersection-related crashes in police-reported narratives. Using 2019 crash data from the Iowa Department of Transportation, we implemented and compared a comprehensive set of models including Support Vector Machine (SVM), XGBoost, BERT Sentence Embeddings, BERT Word Embeddings, and Albert Model. Model performance was systematically validated against expert reviews of potentially misclassified narratives, providing a rigorous assessment of classification accuracy. Results demonstrated that while traditional ML methods exhibited superior overall performance compared to some DL approaches, the Albert Model achieved the highest agreement with expert classifications (73% with Expert 1) and original tabular data (58%) . Statistical analysis revealed that the Albert Model maintained performance levels similar to inter-expert consistency rates, significantly outperforming other approaches particularly on ambiguous narratives. This work addresses a critical gap in transportation safety research through multi-modal integration analysis, which achieved a 54.2% reduction in error rates by combining narrative text with structured crash data. We conclude that hybrid approaches combining automated classification with targeted expert review offer a practical methodology for improving  \n0 *These authors contributed equally to this work.  \ncrash data quality, with substantial implications for transportation safety management and policy development.  \nKeywords: Misclassification, crash narratives, machine learning, deep learning, natural language processing, expert opinion  \n1 Introduction  \nPolice-reported crash data are fundamental to transportation safety management, informing activities ranging from crash prediction to countermeasure evaluation and policy development (AASHTO, 2010) . The accuracy of these data critically influences the reliability of subsequent safety analyses, yet data quality issues—particularly misclassification of key variables—remain prevalent and understudied (Montella et al. , 2013; Abay, 2015; Pasindu, 2019) . When crash features such as intersection involvement are incorrectly coded, this can significantly impact safety program effectiveness and resource allocation decisions (Pasindu, 2019; Abdulhafedh, 2017) .  \nCrash reports typically include both structured data elements and narrative descriptions. While structured data provide quantifiable information, crash narratives offer detailed accounts that may contain critical contextual information about crash circumstances (Kim et al., 2021; Trueblood et al., 2019) . These narratives represent a valuable but underutilized resource that could potentially identify and correct misclassifications in the structured data. Previous research has demonstrated the utility of crash narratives for frequency analysis (Boggs et al., 2020), severity analysis (Arteaga et al., 2020), network screening (Ambros Jiriet al., 2016), and countermeasure selection (Saha, 2022) . However, the manual review of narratives is time-intensive and impractical at scale, creating a need for automated methods (McCullough and Smith, 1998; Williamson et al., 2001) .  \nT","cbCaikoN78XKvipz","https://ap.wps.com/l/cbCaikoN78XKvipz","pdf",452907,6,1,32,"English","en",105,"# Abstract\n# Introduction\n## Crash data quality and misclassification\n## Role of crash narratives\n## Limitations of manual narrative review\n## Need for automated ML/DL approaches","[{\"question\":\"What problem does the study address in transportation safety data?\",\"answer\":\"It addresses the misclassification of intersection-related crash variables in police-reported data, which can undermine safety analyses and decision-making.\"},{\"question\":\"Which models are compared for detecting potentially misclassified narratives?\",\"answer\":\"The study compares SVM, XGBoost, BERT sentence embeddings, BERT word embeddings, and the Albert model using 2019 Iowa crash data.\"},{\"question\":\"How are model results validated in the research?\",\"answer\":\"Model performance is systematically validated against expert reviews of potentially misclassified crash narratives, and also compared with original tabular data.\"}]","Identification of Potentially Misclassified Crash Narratives using Machine Learning (ML) and Deep Learning (DL) | PDF",1785904383,81,{"code":4,"msg":32,"data":33},"ok",{"site_id":25,"language":24,"slug":34,"title":13,"keywords":35,"description":14,"schema_data":36,"social_meta":88,"head_meta":90,"extra_data":92,"updated_unix":29},"identification-of-potentially-misclassified-crash-narratives-using-machine-learning-ml-and-deep-learning-dl","",{"@graph":37,"@context":87},[38,55,70],{"@type":39,"itemListElement":40},"BreadcrumbList",[41,45,49,52],{"item":42,"name":43,"@type":44,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":46,"name":47,"@type":44,"position":48},"https://docshare.wps.com/document/","Document",2,{"item":50,"name":12,"@type":44,"position":51},"https://docshare.wps.com/document/research-report/",3,{"item":53,"name":13,"@type":44,"position":54},"https://docshare.wps.com/document/identification-of-potentially-misclassified-crash-narratives-using-machine-learning-ml-and-deep-learning-dl/126309/",4,{"url":53,"name":13,"@type":56,"author":57,"headline":13,"publisher":59,"fileFormat":62,"inLanguage":24,"description":14,"dateModified":63,"datePublished":64,"encodingFormat":62,"isAccessibleForFree":65,"interactionStatistic":66},"DigitalDocument",{"name":9,"@type":58},"Person",{"url":42,"name":60,"@type":61},"DocShare","Organization","application/pdf","2026-08-22","2026-08-05",true,{"@type":67,"interactionType":68,"userInteractionCount":20},"InteractionCounter",{"@type":69},"ViewAction",{"@type":71,"mainEntity":72},"FAQPage",[73,79,83],{"name":74,"@type":75,"acceptedAnswer":76},"What problem does the study address in transportation safety data?","Question",{"text":77,"@type":78},"It addresses the misclassification of intersection-related crash variables in police-reported data, which can undermine safety analyses and decision-making.","Answer",{"name":80,"@type":75,"acceptedAnswer":81},"Which models are compared for detecting potentially misclassified narratives?",{"text":82,"@type":78},"The study compares SVM, XGBoost, BERT sentence embeddings, BERT word embeddings, and the Albert model using 2019 Iowa crash data.",{"name":84,"@type":75,"acceptedAnswer":85},"How are model results validated in the research?",{"text":86,"@type":78},"Model performance is systematically validated against expert reviews of potentially misclassified crash narratives, and also compared with original tabular data.","https://schema.org",{"og:url":53,"og:type":89,"og:title":13,"og:site_name":60,"og:description":14},"article",{"robots":91,"canonical":53},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":94},[95,99,103,107,112,116,121,124,129,132,136],{"id":21,"doc_module":4,"doc_module_name":47,"category_name":96,"show_sort_weight":97,"slug":98},"Story & Novel",90,"story-novel",{"id":48,"doc_module":4,"doc_module_name":47,"category_name":100,"show_sort_weight":101,"slug":102},"Literature",80,"literature",{"id":54,"doc_module":4,"doc_module_name":47,"category_name":104,"show_sort_weight":105,"slug":106},"Exam",70,"exam",{"id":108,"doc_module":4,"doc_module_name":47,"category_name":109,"show_sort_weight":110,"slug":111},5,"Comic",60,"comic",{"id":20,"doc_module":4,"doc_module_name":47,"category_name":113,"show_sort_weight":114,"slug":115},"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":47,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":47,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":47,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":47,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":47,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":47,"category_name":138,"show_sort_weight":108,"slug":139},19,"General","general"]