[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-125678-en":3,"doc-seo-125678-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},125678,8796095360427,"Lucas Martin","https://ap-avatar.wpscdn.com/davatar_994ba38a5ba835b3df7d355c54d3ed8d",8,"Research & Report","An automated machine learning approach for classifying infrastructure cost data","Infrastructure project cost data are frequently unstructured and inconsistent, making cross-organization comparison and benchmarking hard to achieve using manual classification. The described work automates this process by using natural language processing to extract relevant keywords from text descriptions and applying machine learning classifiers to replicate expert judgments. It targets “extra over” cost items, conversion factors, and correct work breakdown structure (WBS) categories. Results report 94% correct “extra over” classification, 90% conversion predictions with 87% conversion-factor accuracy, and 72% WBS category accuracy, enabling faster and more reliable cost structuring.","DOI: 10.1111/mice.13114  \nINDUSTRIAL APPLICATION  \nAn automated machine learning approach for classifying infrastructure cost data  \nDaniel Adanza Dopazo1  Lamine Mahdjoubi1  Bill Gething1   \nAbdul-Majeed Mahamadu2  \n1 School of architectue and environment, University of the West England, Bristol, England, UK  \n2 Department of architecture, University of College London, London, England, UK  \nCorrespondence  \nDaniel Adanza Dopazo, University of the West England, Bristol, School of architecture and environment, Coldharbour Ln, Stoke Gifford, Bristol BS16 1QY, UK. [Email:](Email: Daniel.Dopazo@uwe.ac.uk)[ Daniel.Dopazo@uwe.ac.uk](Email: Daniel.Dopazo@uwe.ac.uk)  \nFunding information  \nInnovate UK, Grant/Award Number: 45382; Transport Infrastructure Efficiency Strategy Living Lab, Grant/Award Number: 45382  \nAbstract  \nData on infrastructure project costs are often unstructured and lack consistency. To enable costs to be compared within and between organizations, large amounts of data must be classified to a common standard, typically a manual process. This is time-consuming, error-prone, inconsistent, and subjective, as it is based on human judgment. This paper describes a novel approach for automating the process by harnessing natural language processing identifying the relevant keywords in the text descriptions and implementing machine learning classifiersto emulate the expert’s knowledge. The task was to identify “extra over” cost items, conversion factors, and to recognize the correct work breakdown structure (WBS) category. The results show that 94% of the “extra over” cases were correctly classified, and 90% of cases that needed conversion, correctly predicting an associated conversion factor with 87% accuracy. Finally, the WBS categories were identified with 72% accuracy. The approach has the potential to provide a step change in the speed and accuracy of structuring and classifying infrastructure cost data for benchmarking.  \n1  INTRODUCTION  \nThe increasingly bigger amount of data in infrastructure project cost, its inconsistencies in structure, and the variety of formats, even within a single organization, makes the comparison process, the analysis, and the decision-making tasks difficult and unreliable.  \nThe problem does not have an easy solution, and extracting and reclassifying this information into a common standard involves the manipulation of large amounts of information systematically (Ahiaga-Dagbui & Smith, 2012) . This task is still largely performed manually by many construction companies. Due to the size of the data involved, this task is time-consuming, error-prone, and  \nmay also be inconsistent and subjective since it is based on human judgment (Martínez-Rojas et al., 2015) .  \nAdditionally, processing large amounts of often disparate cost data present additional challenges such as: including varying formats and unstructured content of cost documents and lack of standardization of the data schemas, as well as poor cost classification practices across the industry (Matthews et al., 2022; NIC, 2021) .  \nThis, in turn, results in severe challenges in data analytics and benchmarking, within and between organizations let alone across the wider infrastructure industry. In fact, according to the Data Management Association, organizations spend between 10% and 30% of their revenue on managing data quality issues (GDQH, 2021) . This includes  \nThis is an open access article under the terms of the Creative Commons Attribution License, which permits use, distribution and reproduction in any medium, provided the original work is properly cited.  \n© 2023 The Authors. Computer-Aided Civil and Infrastructure Engineering published by Wiley Periodicals LLC on behalf of Editor.  \nDOPAZO et al.  \ncosts associated with pre-processing and the general loss of information due to inconsistency and variable quality.  \nA recent report confirmed that 32% of construction project data is often of poor quality, being inaccurate, unusa","cbCaivwaiG9c0dvZ","https://ap.wps.com/l/cbCaivwaiG9c0dvZ","pdf",881926,1,16,"English","en",105,"# Abstract\n# Introduction","[{\"question\":\"Why is manual classification of infrastructure cost data difficult?\",\"answer\":\"Cost descriptions are often unstructured and inconsistent across organizations, so manual work is time-consuming, error-prone, and relies on subjective human judgment.\"},{\"question\":\"What does the automated approach extract from cost text descriptions?\",\"answer\":\"It uses natural language processing to identify relevant keywords, then applies machine learning classifiers to identify “extra over” items, conversion factors, and the correct WBS category.\"},{\"question\":\"How accurate is the proposed method for the different classification tasks?\",\"answer\":\"It correctly classifies 94% of “extra over” cases, predicts conversion-needed cases with 90% accuracy while achieving 87% accuracy for conversion factors, and identifies WBS categories with 72% accuracy.\"}]","An automated machine learning approach for classifying infrastructure cost data | PDF",1785900612,40,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"an-automated-machine-learning-approach-for-classifying-infrastructure-cost-data","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/an-automated-machine-learning-approach-for-classifying-infrastructure-cost-data/125678/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-05",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why is manual classification of infrastructure cost data difficult?","Question",{"text":75,"@type":76},"Cost descriptions are often unstructured and inconsistent across organizations, so manual work is time-consuming, error-prone, and relies on subjective human judgment.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What does the automated approach extract from cost text descriptions?",{"text":80,"@type":76},"It uses natural language processing to identify relevant keywords, then applies machine learning classifiers to identify “extra over” items, conversion factors, and the correct WBS category.",{"name":82,"@type":73,"acceptedAnswer":83},"How accurate is the proposed method for the different classification tasks?",{"text":84,"@type":76},"It correctly classifies 94% of “extra over” cases, predicts conversion-needed cases with 90% accuracy while achieving 87% accuracy for conversion factors, and identifies WBS categories with 72% accuracy.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,119,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":29,"slug":118},7,"Healthcare","healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]