[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-127934-en":3,"doc-seo-127934-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},127934,687207024478,"Liam","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","A Comprehensive Comparison of Missing Data Procedures for Tree-Based Machine Learning Methods - Dissertation Abstract","Missingness must be addressed before tree-based machine learning models such as decision trees, random forests, and QUalitative INteraction Trees (QUINT) can generate reliable predictions. Existing comparisons of missing data procedures for tree models are fragmented and often examine only limited conditions or methods, and QUINT-specific applications have received little prior study. This dissertation synthesizes the literature, reviews psychological applications, and compares popular modern missing data methods for QUINT via simulation.","UCLA  \nUCLA Electronic Theses and Dissertations  \nTitle  \nA Comprehensive Comparison of Missing Data Procedures for Tree-Based Machine Learning Methods  \nPermalink  \n[https://escholarship.org/uc/item/4ws2h9dc](https://escholarship.org/uc/item/4ws2h9dc)  \nAuthor  \nTibbe, Tristan Dale  \nPublication Date  \n2024  \nPeer reviewed|Thesis/dissertation  \n[eScholarship.org](eScholarship.org) Powered by the California Digital Library  \nUniversity of California  \nUNIVERSITY OF CALIFORNIA Los Angeles  \nA Comprehensive Comparison of Missing Data Procedures for Tree-Based Machine Learning Methods  \nA dissertation submitted in partial satisfaction of the requirements for the degree Doctor of Philosophy in Psychology  \nby  \nTristan Dale Tibbe  \n2024  \n© Copyright by Tristan Dale Tibbe 2024  \nABSTRACT OF THE DISSERTATION  \nA Comprehensive Comparison of  \nMissing Data Procedures for  \nTree-Based Machine Learning Methods  \nby  \nTristan Dale Tibbe  \nDoctor of Philosophy in Psychology  \nUniversity of California, Los Angeles, 2024  \nProfessor Amanda K. Montoya, Chair  \nAs machine learning techniques increase in popularity among psychology researchers, decision trees—and their offshoots such as random forests and qualitative interaction trees (QUINT)—have received special attention due to their interpretability and ease of use. Although these tree-based methods are versatile in terms of the types and number of relationships they can model, missingness still needs to be addressed before they can produce predictions. The current literature comparing missing data methods available for tree-based models is fragmented, with only subsets of conditions or methods examined in each study. Furthermore, the application of missing data methods to the specialized tree-based algorithm QUINT has largely gone uninvestigated in previous research. Thus, there is a great need to clarify which missing data methods are best to use in which situations, especially with niche tree-based models like QUINT. Since tree-based methods are quick to set up and easy to interpret, it is particularly important that users who may not be experienced in machine learning or statistics receive guidance on how they should manage missingness in their data before they apply such methods. In order to consolidate the research that has already been  \ndone on missing data methods with tree-based models, this dissertation provides a summary of the existing literature on the topic, introducing the methods and factors that are important to consider when choosing how to deal with missingness in datasets for tree models. Also, to understand which missing data methods are currently applied to tree-based models in psychological studies and under which conditions, recent substantive research articles were reviewed and information about their datasets and methodologies were recorded. Finally, to extend the knowledge accumulated in the literature, a plethora of popular/modern missing data methods for tree models were applied to the QUINT algorithm and compared in a simulation study, varying factors such as the amount of missingness, the type of missingness, and where the missingness appeared in the data to cover a variety of scenarios that may arise in the real world. The results reveal that, in terms of both prediction accuracy and variable selection, imputation methods—specifically regression and hot deck imputation—and missingness incorporated in attributes (MIA) are able to address missing data and produce QUINT models of consistently high quality. The complete case method, on the other hand, should be avoided due to its highly variable performance and inconsistent nature, leading to models that differ greatly from would have been produced had the data been complete.  \nThis dissertation of Tristan Dale Tibbe is approved.  \nCraig Kyle Enders  \nHan Du  \nChristina Michelle Ramirez Amanda K. Montoya, Committee Chair  \nUniversity of California, Los Angeles  \n2024  \nDEDICATION  \nThis dissertation is dedicat","cbCais8qIfISnWGq","https://ap.wps.com/l/cbCais8qIfISnWGq","pdf",8829314,1,151,"English","en",105,"# Introduction\n## Decision Trees\n## Random Forests\n## QUalitative INteraction Trees (QUINT)\n## Missing Data\n## Missing Data Procedures\n## Missing Data Method Recommendations\n# Meta-Scientific Review\n## Meta-Scientific Review Methods\n## Meta-Scientific Review Results\n## Sample Information\n## Variable Information\n## Methodology Information\n# Simulation\n## Manipulated Factors\n## Complete Datasets\n## Missing Data Generation\n## Missing Data Methods\n## Measured Outcomes\n## Simulation Results","[{\"question\":\"Why is handling missing data necessary for tree-based machine learning methods?\",\"answer\":\"Tree-based methods need missingness handled before they can produce predictions. Without appropriate procedures, model performance and variable selection can become unreliable.\"},{\"question\":\"What gaps exist in prior research on missing data methods for tree models?\",\"answer\":\"Prior literature is fragmented, with studies often covering only subsets of conditions or methods. Additionally, missing data methods applied to the QUINT algorithm have been largely unexplored.\"},{\"question\":\"Which missing data approaches performed best for QUINT in the dissertation’s simulation study?\",\"answer\":\"Imputation methods—especially regression and hot deck imputation—and missingness incorporated in attributes (MIA) produced consistently high-quality QUINT models. The complete case method was less reliable and should be avoided.\"}]","A Comprehensive Comparison of Missing Data Procedures for Tree-Based Machine Learning Methods - Dissertation Abstract | PDF",1785943076,381,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"a-comprehensive-comparison-of-missing-data-procedures-for-tree-based-machine-learning-methods-dissertation-abstract","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/a-comprehensive-comparison-of-missing-data-procedures-for-tree-based-machine-learning-methods-dissertation-abstract/127934/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-25","2026-08-05",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"Why is handling missing data necessary for tree-based machine learning methods?","Question",{"text":76,"@type":77},"Tree-based methods need missingness handled before they can produce predictions. Without appropriate procedures, model performance and variable selection can become unreliable.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"What gaps exist in prior research on missing data methods for tree models?",{"text":81,"@type":77},"Prior literature is fragmented, with studies often covering only subsets of conditions or methods. Additionally, missing data methods applied to the QUINT algorithm have been largely unexplored.",{"name":83,"@type":74,"acceptedAnswer":84},"Which missing data approaches performed best for QUINT in the dissertation’s simulation study?",{"text":85,"@type":77},"Imputation methods—especially regression and hot deck imputation—and missingness incorporated in attributes (MIA) produced consistently high-quality QUINT models. The complete case method was less reliable and should be avoided.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":46,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":46,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]