[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-121874-en":3,"doc-seo-121874-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},121874,4810365810221,"Aurora","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","Incorporating Machine Learning into Sociological Model-Building - Mark D. Verhagen","Quantitative sociologists often rely on simple linear functional forms, yet limited guidance exists on whether these forms match the underlying data-generating process. This work proposes a framework that uses flexible machine learning methods to estimate fit potential using the same covariates as the researcher’s hypothesized model. When ML fit potential substantially exceeds the hypothesized form, the approach signals missing complexity and motivates model refinement. Explainable AI tools such as Shapley values are used to interpret ML and improve the functional form, with demonstrations via simulation and real-world cases.","Standard Length Article  \nIncorporating Machine Learning into Sociological Model-Building  \nMark D. Verhagen 1,2,3  \nSociological Methodology 1–52  \n􀀂 The Author(s) 2024  \nArticle reuse guidelines: DOI: 10.1177/00811750231217734  \n[http://sm.sagepub.com](http://sm.sagepub.com)  \nAbstract  \nQuantitative sociologists frequently use simple linear functional forms to estimate associations among variables. However, there is little guidance on whether such simple functional forms correctly reflect the underlying data-generating process. Incorrect model specification can lead to misspecification bias, and a lack of scrutiny of functional forms fosters interference of researcher degrees of freedom in sociological work. In this article, I propose a framework that uses flexible machine learning (ML) methods to provide an indication of the fit potential in a dataset containing the exact same covariates as a researcher’s hypothesized model. When this ML-based fit potential strongly outperforms the researcher’s selfhypothesized functional form, it implies a lack of complexity in the latter. Advances in the field of explainable AI, like the increasingly popular Shapley values, can be used to generate understanding into the ML model such that the researcher’s original functional form can be improved accordingly. The proposed framework aims to use ML beyond solely predictive questions, helping sociologists exploit the potential of ML to identify intricate patterns in data to specify better-fitting, interpretable models. I illustrate the proposed framework using a simulation and real-world examples.  \nKeywords  \nMachine learning, Misspecification, Explainable A.I., Computational methods  \nIt is common knowledge that valid inference crucially depends on a correctly specified relationship between the outcome of interest, y, and the explanatory variables, X (Buja, Brown, et al. 2019; Cameron and Trivedi 2005; Long and Trivedi 1992) . In practice, much more attention is typically paid to identifying relevant variables to include in a model rather than to making sure the functional relationship among these variables is correctly specified. This is evidenced by the fact that variables are often simply assumed to affect the outcome in a linear and additive way (Hindman 2015) . However, there is little reason to believe linear additive models appropriately reflect the underlying data-generating process (DGP) . At the same time, estimating incorrectly specified models can lead to biased findings; there are noteworthy examples throughout the social sciences—and likely many more that have gone unnoticed—where more complicated functional relationships, which might include nonlinearities  \n1Leverhulme Centre for Demographic Science, Oxford, UK 2Nuffield College, University of Oxford, Oxford, UK  \n3Department of Sociology, University of Oxford, Oxford, UK Corresponding Author:  \nMark D. Verhagen, University of Oxford, New Rd, Oxford, OX1 2JD, UK. [Email: mark.verhagen@nuffield.ox.ac.uk](Email: mark.verhagen@nuffield.ox.ac.uk)  \nor interactions, have led to reversed findings (Christensen and Christensen 2014; Dougherty et al. 2015; Freedman 2009; Heckman, Humphries, and Veramendi 2018; McClintock 2017; Mu˜noz and Young 2018) . Despite this type of criticism of the standard linear additive model having been around for many decades, it remains the workhorse throughout most empirical sociology today (Abbott 1988; Berk 2004; Duncan 1984; Lundberg, Johnson, and Stewart 2021) . In this article, I incorporate methods from machine learning (ML) and explainable A.I. (X-AI) into the standard empirical workflow to help sociologists (1) assess whether their hypothesized model fits the data well by comparing its fit against a flexible ML model, and (2) improve their model when it does not accurately represent the patterns in the data by unpacking the ML model using X-AI techniques.  \nIn the past, simple models like the linear additive functional form were often a necessi","cbCairmokmwQjA0G","https://ap.wps.com/l/cbCairmokmwQjA0G","pdf",4064304,1,52,"English","en",105,"# Abstract\n# Core problem and motivation\n## Limits of linear additive functional forms\n## Consequences of misspecification bias\n# Proposed ML + explainable AI framework\n## Compare fit potential against the hypothesized model\n## Use explainable AI to improve interpretability\n# Implementation in the quantitative workflow\n## Ensemble ML estimation (Super Learner)\n## Interpret and revise functional form\n# Evaluation\n## Simulation and real-world examples","[{\"question\":\"Why do quantitative sociologists need help with functional-form specification?\",\"answer\":\"Researchers often focus on selecting variables, while functional relationships are frequently assumed linear and additive. Incorrect functional forms can produce misspecification bias and misleading findings.\"},{\"question\":\"How does the framework use machine learning to assess a hypothesized sociological model?\",\"answer\":\"It estimates the researcher’s hypothesized function and, in parallel, a flexible ML model using the same covariates and inferential logic. ML fit potential is then compared to reveal whether the hypothesized form lacks complexity.\"},{\"question\":\"How can explainable AI improve the hypothesized functional form?\",\"answer\":\"Explainable AI methods such as Shapley values help interpret the ML model. Insights from this interpretation guide improvements so the original functional form better matches the patterns in the data.\"}]","Incorporating Machine Learning into Sociological Model-Building - Mark D. Verhagen | PDF",1785807373,131,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"incorporating-machine-learning-into-sociological-model-building-mark-d-verhagen","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/incorporating-machine-learning-into-sociological-model-building-mark-d-verhagen/121874/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-05","2026-08-04",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"Why do quantitative sociologists need help with functional-form specification?","Question",{"text":76,"@type":77},"Researchers often focus on selecting variables, while functional relationships are frequently assumed linear and additive. Incorrect functional forms can produce misspecification bias and misleading findings.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does the framework use machine learning to assess a hypothesized sociological model?",{"text":81,"@type":77},"It estimates the researcher’s hypothesized function and, in parallel, a flexible ML model using the same covariates and inferential logic. ML fit potential is then compared to reveal whether the hypothesized form lacks complexity.",{"name":83,"@type":74,"acceptedAnswer":84},"How can explainable AI improve the hypothesized functional form?",{"text":85,"@type":77},"Explainable AI methods such as Shapley values help interpret the ML model. Insights from this interpretation guide improvements so the original functional form better matches the patterns in the data.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":46,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":46,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]