[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-121037-en":3,"doc-seo-121037-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},121037,2336464648322,"Aria","https://ap-avatar.wpscdn.com/avatar/2200025388227c56fec?_k=1778556882303663488",8,"Research & Report","Building a Data Pipeline and Machine Learning Model for Insurance Data - Honors Thesis","Insurance telematics combines GPS tracking, computational analytics, data processing, and machine learning to support insurers with better underwriting decisions. National Indemnity sought to add a telematics component to its underwriting workflow and commissioned an undergraduate honors thesis project. The work reviews relevant prior solutions and then designs a data pipeline and machine learning model to generate predictions of drivers’ risk levels. The resulting approach enables more accurate insurance rate optimization by reflecting insured risk more effectively.","University of Nebraska-Lincoln  \nDigitalCommons@University of Nebraska-Lincoln  \n\n| Honors Theses | Honors Program |\n| --- | --- |\n| Spring 5-13-2024\u003Cbr>Building a Data Pipeline and Machine Learning Model for Insurance Data\u003Cbr>Connor Weyers\u003Cbr>University of Nebraska-Lincoln\u003Cbr>Follow this and additional works at: [https://digitalcommons.unl.edu/honorstheses](https://digitalcommons.unl.edu/honorstheses)\u003Cbr> Part of the Databases and Information Systems Commons, Gifted Education Commons, Higher Education Commons, Other Computer Sciences Commons, and the Other Education Commons |  |\n\nWeyers, Connor, \"Building a Data Pipeline and Machine Learning Model for Insurance Data\" (2024) . Honors Theses. 729.  \n[https://digitalcommons.unl.edu/honorstheses/729](https://digitalcommons.unl.edu/honorstheses/729)  \nThis Thesis is brought to you for free and open access by the Honors Program at DigitalCommons@University of Nebraska-Lincoln. It has been accepted for inclusion in Honors Theses by an authorized administrator of DigitalCommons@University of Nebraska-Lincoln.  \nUniversity of Nebraska-Lincoln  \nUndergraudate Thesis for University Honors  \nProgram  \nBuilding a Data Pipeline and Machine Learning Model for Insurance Data  \nConnor Weyers, BS Computer Science School of Computing  \nSupervised by  \nProf. Vinodchandran Variyam  \nMay 13, 2024  \nContents  \n1 Introduction .................................. 1  \n1.1 Challenges for Machine Learning .................. 1  \n2 Research .................................... 1  \n2.1 Data Ingestion ............................ 1  \n2.2 Data Storage ............................. 2  \n2.3 Data Processing ............................ 3  \n2.4 Machine Learning Models ...................... 5  \n3 Data Pipeline Implementation ........................ 8  \n3.1 Data Ingestion ............................ 9  \n3.2 Data Storage ............................. 9  \n3.3 Data Processing ............................ 10  \n4 Machine Learning Model Implementation .................. 11  \n5 Conclusion ................................... 13  \nAbstract  \nInsurance telematics is an emerging and exciting field. It combines the advancementsin GPS tracking, computational analytics, data processing, and machine learning into a useful tool to help insurance companies make the best product for their consumers. This is why National Indemnity looked to implement a telematics portion to their business processes of underwriting insurance policies and sponsored a School of Computing Senior Design project. In this report, we will first review existing solutions that been used to solve problems and subproblems similar to that we are given in this project. We then propose designs for the data pipeline and machine learning model that will optimal in providing predictions on the risk level of drivers. National Indemnity will be able to use this project to leverage predictions in order to optimize insurance rates to more accurately account for risk among the insured.  \nKeywords: Computer Science, Machine Learning, Insurance, Telematics, Data Processing, Data Analytics  \n1 Introduction  \nIn the last twenty years there has been a fast development in the technology being used for all types of businesses. This had led to ever-increasing competition in the market to stay up to date with industry standards. The areas that have particularly revolutionized by technology include data gathering, data processing, and data analysis. Recently, machine learning had emerged as an important and popular data analysis method. Abroad goal of machine learning is to algorithmically discover hidden trends in datasets. The improvements in machine learning models and computational efficiency saw machine learning being used in a variety of contexts. Machine learning began to be applied to the insurance industry with increasing frequency, including in this project for National Indemnity Company. But for machine learning to be best applied, it must have quality data. For this da","cbCaifaXDzM4IV31","https://ap.wps.com/l/cbCaifaXDzM4IV31","pdf",179185,1,20,"English","en",105,"# Introduction\n## Challenges for Machine Learning\n# Research\n## Data Ingestion\n## Data Storage\n## Data Processing\n## Machine Learning Models\n# Data Pipeline Implementation\n## Data Ingestion\n## Data Storage\n## Data Processing\n# Machine Learning Model Implementation\n# Conclusion","[{\"question\":\"What problem does the thesis address in the insurance industry?\",\"answer\":\"It addresses how to use telematics data and machine learning to predict driver risk levels for more accurate underwriting decisions and insurance pricing.\"},{\"question\":\"How does the project ensure the machine learning model can make useful predictions?\",\"answer\":\"It focuses on building a data pipeline that ingests data, stores it, and processes it into uniform form, then trains and evaluates models to maintain accuracy.\"},{\"question\":\"Why is the data pipeline described as needing flexibility?\",\"answer\":\"Insurance data and its sources update over time, so the pipeline and model must be highly adaptable to changing data structures and new technologies to avoid performance degradation.\"}]","Building a Data Pipeline and Machine Learning Model for Insurance Data - Honors Thesis | PDF",1785733426,50,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"building-a-data-pipeline-and-machine-learning-model-for-insurance-data-honors-thesis","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/building-a-data-pipeline-and-machine-learning-model-for-insurance-data-honors-thesis/121037/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the thesis address in the insurance industry?","Question",{"text":75,"@type":76},"It addresses how to use telematics data and machine learning to predict driver risk levels for more accurate underwriting decisions and insurance pricing.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the project ensure the machine learning model can make useful predictions?",{"text":80,"@type":76},"It focuses on building a data pipeline that ingests data, stores it, and processes it into uniform form, then trains and evaluates models to maintain accuracy.",{"name":82,"@type":73,"acceptedAnswer":83},"Why is the data pipeline described as needing flexibility?",{"text":84,"@type":76},"Insurance data and its sources update over time, so the pipeline and model must be highly adaptable to changing data structures and new technologies to avoid performance degradation.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,114,119,122,126,129,133],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":29,"slug":113},6,"Technology","technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":21,"slug":125},9,"Religion & Spirituality","religion-spirituality",{"id":21,"doc_module":4,"doc_module_name":46,"category_name":127,"show_sort_weight":21,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":46,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":46,"category_name":135,"show_sort_weight":106,"slug":136},19,"General","general"]