[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-118810-en":3,"doc-seo-118810-105":29,"detail-sidebar-cat-0-en-105":90},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":11,"language":21,"language_code":22,"site_id":23,"html_lang":22,"table_of_contents":24,"faqs":25,"seo_title":26,"seo_description":14,"update_tm":27,"read_time":28},118810,962075006959,"Anda","https://ap-avatar.wpscdn.com/avatar/e0002397efbe92a78e?_k=1776741047341049297",8,"Research & Report","Subjectivity in Unsupervised Machine Learning Model Selection","Model selection is a necessary step in unsupervised machine learning, yet it remains inherently subjective despite many available criteria and metrics. High subjectivity can undermine repeatability and reproducibility and cast doubt on whether deployed models will be robust in real-world settings. This study examines how modelers’ preferences shape outcomes by using a Hidden Markov Model example, asking 33 participants and three large language models to choose models across three scenarios. Results show inconsistent selections, particularly when criteria disagree, and trace subjectivity to differing views on criteria importance, parsimony, and dataset size. The findings motivate more standardized documentation of these choices.","Subjectivity in Unsupervised Machine Learning Model Selection  \nWanyi Chen 1 , Mary L. Cummings 2  \n1Duke University  \n2 George Mason University  \n[wc151@duke.edu](wc151@duke.edu), [cummings@gmu.edu](cummings@gmu.edu)  \narXiv :2309 .0020 1v 1 [ cs .LG] 1 Sep 2023  \nAbstract  \nModel selection is a necessary step in unsupervised machine learning. Despite numerous criteria and metrics, model selection remains subjective. A high degree of subjectivity may lead to questions about repeatability and reproducibility of various machine learning studies and doubts about the robustness of models deployed in the real world. Yet, the impact of modelers’ preferences on model selection outcomes remains largely unexplored. This study uses the Hidden Markov Model as an example to investigate the subjectivity involved in model selection. We asked 33 participants and three Large Language Models (LLMs) to make model selections in three scenarios. Results revealed variability and inconsistencies in both the participants’ and the LLMs’ choices, especially when different criteria and metrics disagree. Sources of subjectivity include varying opinions on the importance of different criteria and metrics, differing views on how parsimonious a model should be, and how the size of a dataset should influence model selection. The results underscore the importance of developing a more standardized way to document subjective choices made in model selection processes.  \n1 Introduction  \nIn a world of abundant data, unsupervised machine learning (ML) is popular for discovering patterns and structures in data without needing labels. Selecting the best model is a necessary step, as even slightly different models can lead to different interpretations and decisions. For instance, psychologists can use unsupervised ML to uncover patterns inhuman learning behaviors (Visser, Raijmakers, and Molenaar 2002) . Different interpretations of such models would likely lead to different training designs, with some less optimal than others.  \nThere are a few objectives in selecting the best model in unsupervised ML. On the one hand, a model should accurately describe the data. On the other, a parsimonious model is often more desirable because it is more interpretable and less likely to overfit. However, trade-offs exist between accuracy and parsimony. A model with more variables is more likely accurate but less parsimonious (Dziak et al. 2020) .  \nSeveral criteria guide model selection, including information criteria (IC) such as Akaike Information Criterion (AIC)(Akaike 1974) and Bayesian Information Criterion (BIC)(Schwarz 1978) . Cross-validation metrics such as the most  \nconsistent, best worst case, and best average may provide additional guidance.  \nDespite numerous seemingly objective criteria and metrics, model selection remains subjective. Different ICs place varying emphasis on parsimony versus predictive accuracy (Dziak et al. 2020) . ML researchers and practitioners hold various beliefs about the importance of certain criteria or metrics, so the definition of the “best” model becomes subjective. While it is well known that biases in datasets cause biases in machine learning models (Buolamwini and Gebru 2018), little is known about how the modelers’ preferences influence model selection results. The subjective decisions made in the model selection process can be considered the modeler’s “degree of freedom.”  \nThe modeler’s degree of freedom yields insights into the repeatability and reproducibility of results, essential in scientific research (Plesser 2018) . If results are not repeatable or reproducible, the research community cannot critically assess the correctness of scientific claims made and conclusions drawn from the results (Plesser 2018) . Consequently, we cannot have confidence in the robustness of various realworld ML models based on these research.  \nThis study investigates the subjectivity involved in modelselection using the Hidden Markov Model (HMM) as","cbCaivEulDz2zCSq","https://ap.wps.com/l/cbCaivEulDz2zCSq","pdf",210724,1,"English","en",105,"# Introduction\n## Background and related work\n## Information criterion","[{\"question\":\"Why is model selection considered subjective in unsupervised machine learning?\",\"answer\":\"Even with multiple information criteria and metrics, different researchers place different emphasis on parsimony versus predictive accuracy. As a result, the definition of the “best” model varies by preference and judgment.\"},{\"question\":\"How does the study test subjectivity in model selection?\",\"answer\":\"The study uses a Hidden Markov Model example and runs an experiment where 33 participants and three large language models select models in three scenarios.\"},{\"question\":\"What factors drive subjectivity according to the results?\",\"answer\":\"Subjectivity stems from differences in opinions about which criteria and metrics matter most, how parsimonious a model should be, and how dataset size should influence model selection.\"}]","Subjectivity in Unsupervised Machine Learning Model Selection | PDF",1785720376,20,{"code":4,"msg":30,"data":31},"ok",{"site_id":23,"language":22,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":85,"head_meta":87,"extra_data":89,"updated_unix":27},"subjectivity-in-unsupervised-machine-learning-model-selection","",{"@graph":35,"@context":84},[36,53,67],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/subjectivity-in-unsupervised-machine-learning-model-selection/118810/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":22,"description":14,"dateModified":61,"datePublished":61,"encodingFormat":60,"isAccessibleForFree":62,"interactionStatistic":63},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":64,"interactionType":65,"userInteractionCount":4},"InteractionCounter",{"@type":66},"ViewAction",{"@type":68,"mainEntity":69},"FAQPage",[70,76,80],{"name":71,"@type":72,"acceptedAnswer":73},"Why is model selection considered subjective in unsupervised machine learning?","Question",{"text":74,"@type":75},"Even with multiple information criteria and metrics, different researchers place different emphasis on parsimony versus predictive accuracy. As a result, the definition of the “best” model varies by preference and judgment.","Answer",{"name":77,"@type":72,"acceptedAnswer":78},"How does the study test subjectivity in model selection?",{"text":79,"@type":75},"The study uses a Hidden Markov Model example and runs an experiment where 33 participants and three large language models select models in three scenarios.",{"name":81,"@type":72,"acceptedAnswer":82},"What factors drive subjectivity according to the results?",{"text":83,"@type":75},"Subjectivity stems from differences in opinions about which criteria and metrics matter most, how parsimonious a model should be, and how dataset size should influence model selection.","https://schema.org",{"og:url":51,"og:type":86,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":88,"canonical":51},"index,follow",{"doc_id":7,"site_id":23},{"code":4,"msg":5,"data":91},[92,96,100,104,109,114,119,122,126,129,133],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":93,"show_sort_weight":94,"slug":95},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":97,"show_sort_weight":98,"slug":99},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":101,"show_sort_weight":102,"slug":103},"Exam",70,"exam",{"id":105,"doc_module":4,"doc_module_name":45,"category_name":106,"show_sort_weight":107,"slug":108},5,"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":45,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":28,"slug":125},9,"Religion & Spirituality","religion-spirituality",{"id":28,"doc_module":4,"doc_module_name":45,"category_name":127,"show_sort_weight":28,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":45,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":45,"category_name":135,"show_sort_weight":105,"slug":136},19,"General","general"]