[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-117217-en":3,"doc-seo-117217-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},117217,5909877438554,"Maeve","https://ap-avatar.wpscdn.com/avatar/5600025385ad2bf12a7?_k=1778553567797529272",8,"Research & Report","Estimating Classification Consistency of Machine Learning Models for Screening Measures","The article introduces quantitative procedures for estimating classification consistency in machine learning models used with screening measures in psychology and medicine. High accuracy is necessary, but effective screening also requires consistent classification, meaning repeated administrations yield the same diagnostic group. The proposed approach models classification inconsistency from sampling error during model fitting and measurement error in item responses using bootstrap and Monte Carlo resampling. Three empirical examples and accompanying R code support applied researchers.","Psychological Assessment  \n2024, Vol. 36, Nos. 6–7, 395–406  \n[https://doi.org/10.1037/pas0001313](https://doi.org/10.1037/pas0001313)  \nEstimating Classiﬁcation Consistency of Machine Learning Models  \nfor Screening Measures  \nOscar Gonzalez 1, A. R. Georgeson2, and William E. Pelham III3  \n1 Department of Psychology and Neuroscience, University of North Carolina at Chapel Hill  \n2 Department of Psychology, Arizona State University  \n3 Department of Psychiatry, University of California San Diego  \nThis article illustrates novel quantitative methods to estimate classiﬁcation consistency in machine learning models  \nused for screening measures. Screening measures are used in psychology and medicine to classify individuals into  \ndiagnostic classiﬁcations. In addition to achieving high accuracy, it is ideal for the screening process to have high  \nclassiﬁcation consistency, which means that respondents would be classiﬁed into the same group every time if the  \nassessment was repeated. Although machine learning models are increasingly being used to predict a screening  \nclassiﬁcation based on individual item responses, methods to describe the classiﬁcation consistency of machine  \nlearning models have not yet been developed. This article addresses this gap by describing methods to estimate  \nclassiﬁcation inconsistency in machine learning models arising from two different sources: sampling error during  \nmodel ﬁtting and measurement error in the item responses. These methods usedataresampling techniques such as  \nthe bootstrap and Monte Carlo sampling. These methods are illustrated using three empirical examples predicting  \na health condition/diagnosis from item responses. R code is provided to facilitate the implementation of the  \nmethods. This article highlights the importance of considering classiﬁcation consistency alongside accuracy when  \nstudying screening measures and provides the tools and guidance necessary for applied researchers to obtain  \nclassiﬁcation consistency indices in their machine learning research on diagnostic assessments.  \nPublic Signiﬁcance Statement  \nRecently, methods for machine learning have been used to predict from a screening measure if  \nindividuals should be ﬂagged for a condition (e.g., as depressed vs. not depressed), but it is unknown if  \nthe models provide consistent screening decisions if respondents were to repeatedly receive the  \nscreening measure. We propose statistical procedures to help researchers determine if a machine  \nlearning model is providing consistent screening decisions.  \nKeywords: screening, machine learning, classiﬁcation consistency, reliability  \nSupplemental materials: [https://doi.org/10.1037/pas0001313.supp](https://doi.org/10.1037/pas0001313.supp)  \nAssessments are typically used in the areas of psychology and medicine to classify individuals. Examples include the Psychopathology Checklist–Revised, which is a measure used to screen for psychopathyin a prison population (Hare, 2003), or the Geriatric Depression Scale, which is a measure used to identify older adults with depression (Sheikh  \n& Yesavage, 1986). In this article, we focus on screening assessments, although the themes and ﬁndings discussed generalize to other classiﬁcation scenarios.  \nIn the screening process, it is important that the decisions about the respondents be accurate and consistent. A decision is accurate  \nOscar Gonzalez  [https://orcid.org/0000-0001-7122-8799](https://orcid.org/0000-0001-7122-8799)  \nA. R. Georgeson  [https://orcid.org/0000-0002-6426-9258](https://orcid.org/0000-0002-6426-9258)  \nWilliam E. Pelham III  [https://orcid.org/0000-0003-1480-570X](https://orcid.org/0000-0003-1480-570X)  \nOscar Gonzalez was supported by the UNC Ann Rankin Cowan Excellence Award for High Impact Research. A. R. Georgeson was supported by the National Institute on Drug Abuse (Grant F32DA053137) . William E. Pelham III was supported by the National Institute on Drug Abuse (Grant DA055935) and the","cbCainhCUzKo3Q3U","https://ap.wps.com/l/cbCainhCUzKo3Q3U","pdf",882611,1,12,"English","en",105,"# Background and Motivation\n# Sources of Classification Inconsistency\n## Sampling Error in Model Fitting\n## Measurement Error in Item Responses\n# Resampling-Based Estimation Methods\n## Bootstrap Techniques\n## Monte Carlo Sampling\n# Empirical Illustrations\n# Implementation Support and Practical Guidance\n# Public Significance Statement","[{\"question\":\"What problem does the article address in machine learning screening models?\",\"answer\":\"It addresses how to estimate classification consistency, not just accuracy, for screening measures that classify individuals into diagnostic groups.\"},{\"question\":\"What are the two sources of classification inconsistency considered?\",\"answer\":\"The methods attribute inconsistency to sampling error during model fitting and measurement error in the item responses.\"},{\"question\":\"How are the proposed indices estimated?\",\"answer\":\"The approach uses resampling techniques such as bootstrap and Monte Carlo sampling, and includes R code to implement the procedures.\"}]","Estimating Classification Consistency of Machine Learning Models for Screening Measures | PDF",1785674455,30,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"estimating-classification-consistency-of-machine-learning-models-for-screening-measures","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/estimating-classification-consistency-of-machine-learning-models-for-screening-measures/117217/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-02",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the article address in machine learning screening models?","Question",{"text":75,"@type":76},"It addresses how to estimate classification consistency, not just accuracy, for screening measures that classify individuals into diagnostic groups.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What are the two sources of classification inconsistency considered?",{"text":80,"@type":76},"The methods attribute inconsistency to sampling error during model fitting and measurement error in the item responses.",{"name":82,"@type":73,"acceptedAnswer":83},"How are the proposed indices estimated?",{"text":84,"@type":76},"The approach uses resampling techniques such as bootstrap and Monte Carlo sampling, and includes R code to implement the procedures.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":29,"slug":121},"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]