[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-120858-en":3,"doc-seo-120858-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},120858,1649267921044,"Ava Thompson","https://us-avatar.wpscdn.com/avatar/1800007509477c92dfb?_k=1782875107921204101",8,"Research & Report","Machine Learning Techniques in Usage-Based Insurance - Use of Telematic Data in Auto Insurance","Big data technologies and in-vehicle devices have accelerated Usage-Based Insurance (UBI) by capturing variables that reflect policyholders’ driving behavior. Collected telematic data, including signals from GPS and sensors, show a strong relationship with accident likelihood, enabling improved risk assessment and more personalized premium setting. This thesis uses a synthetic Canadian insurance dataset to predict accident risk using logistic regression, random forests, gradient boosting trees, and feed-forward neural networks, complemented by Shapley decomposition and feature-removal marginal performance loss for variable importance.","Machine Learning Techniques in Usage-Based Insurance:  \nUse of Telematic Data in Auto Insurance  \nHelia Alipanah  \nA Thesis in  \nThe Department of  \nMathematics and Statistics  \nPresented in Partial Fulfilment of the Requirements for the Degree of Master of Science (Mathematics) at Concordia University  \nMontreal, Quebec, Canada  \nJuly 2023  \n© Helia Alipanah, 2023  \nCONCORDIA UNIVERSITY  \nSchool of Graduate Studies  \nThis is to certify that the thesis prepared By: Helia Alipanah  \nEntitled: Machine Learning Techniques in Usage-Based Insurance:  \nUse of Telematic Data in Auto Insurance  \nand submitted in partial fulfilment of the requirements for the degree of  \nMaster of Science (Mathematics)  \ncomplies with the regulations of the University and meets the accepted standards with respect to originality and quality.  \nSigned by the final Examining Committee:  \n  Examiner  \nDr. M´elina Mailhot  \n  Examiner  \n  Thesis Co-Supervisor  \nDr. J. Garrido  \n  Thesis Co-Supervisor  \nDr. F. Godin  \nApproved by    \nChair of Department or Graduate Program Director  \nDean of Faculty  \nDate    \nAbstract  \nMachine Learning Techniques in Usage-based Insurance: Use of Telematic Data in Auto Insurance  \nby Helia Alipanah  \nThe development of big data technologies and in-vehicle devices has contributed to the growth of Usage-Based Insurance (UBI) in recent years. These in-vehicle devices, such as GPS and sensors, collect certain variables that can represent the driving behaviour of policyholders. This collected data, called telematic data, consist of several variables that have strong relationship with the likelihood of having an accident. Consequently, one can use telematic data to improve risk assessments and personalize car insurance premiums.  \nIn this thesis, a synthetic car insurance dataset emulated from a Canadian-based insurance company is used to investigate the use of telematic data in predicting the likelihood of having an accident. More precisely four machine learning techniques—logistic regression, random forests, gradient boosting trees, and feed-forward neural networks—are employed to predict the risk of having an accident. Actuaries often use white box machine learning methods like logistic regression for risk assessment due to their interpretability. However, these method are unable to detect non-linear relationships between variables accurately. Therefore, more complex machine learning techniques such as random forests, gradient boosting trees, and feed-forward neural networks are used to achieve more accurate risk assessment for accidents.  \nIn addition, two variable importance assessment methods—Shapley decomposition and marginal performance loss upon feature removal—are employed to provide insights into the feature contributions in the overall predictive performance of the models.  \nAcknowledgments  \nI would like to express my heartfelt gratitude to my supervisors, Dr. Jos´e Garrido, and Dr. Fr´ed´eric Godin, for their guidance, support, and invaluable expertise throughout the entire process of this thesis. Their commitment, patience, and insightful feedback have been instrumental in shaping the direction and quality of this research. I am truly grateful for their mentorship and for instilling in me a passion for academic inquiry.  \nI would also like to extend my deepest appreciation to my husband, whose unwavering love, encouragement, and understanding have been my rock during this challenging journey. His belief in me and constant motivation have given me the strength to overcome obstacles and pursue my academic aspirations.  \nFurthermore, I am profoundly grateful to my parents for their endless love, unwavering support, and sacrifices they have made throughout my academic pursuit. Their unwavering belief in my abilities has been a constant source of inspiration and motivation. I am forever indebted to them for instilling in me the values of perseverance, determination, and the importance of education.  \nLastly, I would lik","cbCainS3cQPpUVD4","https://ap.wps.com/l/cbCainS3cQPpUVD4","pdf",2087660,1,64,"English","en",105,"# Introduction\n## Research objectives\n## Limitations\n## Literature review\n## Structure\n# Data\n## Data description\n## Multicollinearity\n## Imbalanced data\n# Methodology\n## Predictive models\n## Performance metrics\n## Results\n## Variable importance assessment\n# Bibliography\n# Appendix A","[{\"question\":\"What is the role of telematic data in usage-based insurance?\",\"answer\":\"Telematic data collected from in-vehicle devices represents driving behavior variables. These variables are used to better estimate accident likelihood and refine risk assessment and premium personalization.\"},{\"question\":\"Which machine learning models are used to predict accident risk?\",\"answer\":\"The thesis applies logistic regression, random forests, gradient boosting trees, and feed-forward neural networks to predict the likelihood of having an accident.\"},{\"question\":\"How does the thesis evaluate which features matter most?\",\"answer\":\"It uses two variable importance assessment methods: Shapley decomposition and marginal performance loss upon feature removal to interpret feature contributions to predictive performance.\"}]","Machine Learning Techniques in Usage-Based Insurance - Use of Telematic Data in Auto Insurance | PDF",1785732373,161,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"machine-learning-techniques-in-usage-based-insurance-use-of-telematic-data-in-auto-insurance","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/machine-learning-techniques-in-usage-based-insurance-use-of-telematic-data-in-auto-insurance/120858/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is the role of telematic data in usage-based insurance?","Question",{"text":75,"@type":76},"Telematic data collected from in-vehicle devices represents driving behavior variables. These variables are used to better estimate accident likelihood and refine risk assessment and premium personalization.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Which machine learning models are used to predict accident risk?",{"text":80,"@type":76},"The thesis applies logistic regression, random forests, gradient boosting trees, and feed-forward neural networks to predict the likelihood of having an accident.",{"name":82,"@type":73,"acceptedAnswer":83},"How does the thesis evaluate which features matter most?",{"text":84,"@type":76},"It uses two variable importance assessment methods: Shapley decomposition and marginal performance loss upon feature removal to interpret feature contributions to predictive performance.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]