[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-120757-en":3,"doc-seo-120757-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},120757,962075006959,"Anda","https://ap-avatar.wpscdn.com/avatar/e0002397efbe92a78e?_k=1776741047341049297",8,"Research & Report","Interpretable Machine Learning for the Social Sciences - Applications in Political Science and Labor Economics","Interpretable Machine Learning for the Social Sciences develops methods that address key limitations of common machine learning approaches for social science research. The thesis targets interpretability for individual predictions and the discovery of latent patterns, rather than only forecasting unseen outcomes. It introduces an explanation approach for black-box sequence models, then applies domain-specific models to political science and labor economics. For political science, it develops a text-based ideal point model combining Bayesian matrix factorization and political theory. For labor economics, it presents transfer learning for career sequences and estimates the history-adjusted gender wage gap using data from labor trajectories.","Interpretable Machine Learning for the Social Sciences: Applications in Political Science and Labor Economics  \nKeyon Vafa  \nSubmitted in partial fulﬁllment of the requirements for the degree of Doctor of Philosophy under the Executive Committee  \nof the Graduate School of Arts and Sciences  \nCOLUMBIA UNIVERSITY  \n© 2023 Keyon Vafa All Rights Reserved  \nAbstract  \nInterpretable Machine Learning for the Social Sciences:  \nApplications in Political Science and Labor Economics  \nKeyon Vafa  \nRecent advances in machine learning offer social scientists a unique opportunity to use  \ndata-driven methods to uncover insights into human behavior. However, current machine learning methods are opaque, ineffective on small social science datasets, and tailored for predicting unseen values rather than estimating parameters from data. In this thesis, we develop interpretable machine learning techniques designed to uncover latent patterns and estimate critical quantities in the social sciences. We focus on two aspects of interpretability: explaining individual model predictions and discovering latent patterns from data. We describe a method for explaining the predictions of general, black-box sequence models. This method approximates a combinatorial objective to elucidate the decision-making processes of sequence models. Next, we narrow our focus to domain-speciﬁc applications. In political science, we develop the text-based ideal point model, a model that quantiﬁes political positions from text. This model marries a classical idea from political science with a Bayesian matrix factorization technique to infer meaningful structure from text. In labor economics, we adapt a model from natural language processing to analyze career trajectories. We describe a transfer learning method that can overcome the constraints posed by small survey datasets. Finally, we adapt this predictive model to estimate an important quantity in labor economics: the history-adjusted gender wage gap.  \nTable of Contents  \nAcknowledgments ........................................ xii  \nDedication ............................................ xiv  \nChapter 1: Introduction .................................... 1  \nChapter 2: Rationales for Sequential Predictions ....................... 4  \n2.1 Introduction ...................................... 4  \n2.2 Sequential Rationales ................................. 6  \n2.3 Greedy Rationalization ................................ 9  \n2.4 Model Compatibility ................................. 10  \n2.4.1 Fine-tuning for Compatibility ......................... 11  \n2.4.2 Compatibility Experiments .......................... 12  \n2.5 Connection to Classiﬁcation Rationales ....................... 13  \n2.6 Related Work ..................................... 15  \n2.7 Experimental Setup .................................. 16  \n2.8 Results and Discussion ................................ 17  \n2.8.1 Language Modeling .............................. 18  \n2.8.2 Machine Translation ............................. 21  \n2.9 Summary ....................................... 23  \nChapter 3: Text-Based Ideal Points .............................. 24  \n3.1 Introduction ...................................... 24  \n3.2 The text-based ideal point model ........................... 26  \n3.2.1 Background: Bayesian ideal points ...................... 27  \n3.2.2 Background: Poisson factorization ...................... 27  \n3.2.3 The text-based ideal point model ....................... 28  \n3.3 Related work ..................................... 31  \n3.4 Inference ........................................ 33  \n3.5 Empirical studies ................................... 34  \n3.5.1 The text-based ideal point model (TBIP) on U.S. Senate speeches ...... 35  \n3.5.2 The TBIP on U.S. Senate tweets ....................... 37  \n3.5.3 Using the TBIP as a descriptive tool ..................... 37  \n3.5.4 2020 Democratic candidates ......................... 40  \n3.6 Summary ................","cbCaidlkwbCmzrPo","https://ap.wps.com/l/cbCaidlkwbCmzrPo","pdf",2386825,1,156,"English","en",105,"# Chapter 1: Introduction\n# Chapter 2: Rationales for Sequential Predictions\n## 2.1 Introduction\n## 2.2 Sequential Rationales\n## 2.3 Greedy Rationalization\n## 2.4 Model Compatibility\n## 2.5 Connection to Classiﬁcation Rationales\n## 2.6 Related Work\n## 2.7 Experimental Setup\n## 2.8 Results and Discussion\n## 2.9 Summary\n# Chapter 3: Text-Based Ideal Points\n## 3.1 Introduction\n## 3.2 The text-based ideal point model\n## 3.3 Related work\n## 3.4 Inference\n## 3.5 Empirical studies\n## 3.6 Summary\n# Chapter 4: CAREER: Transfer Learning for Labor Sequence Data\n## 4.1 Introduction\n## 4.2 CAREER\n## 4.3 Related Work\n## 4.4 Empirical Studies\n## 4.5 Summary\n# Chapter 5: Adjusting the Gender Wage Gap for Full Job History\n## 5.1 Introduction\n## 5.2 Methodology\n## 5.3 Semi-Synthetic Experiments\n## 5.4 Empirical Studies\n## 5.5 Summary\n# Conclusion","[{\"question\":\"What interpretability goals does the thesis focus on?\",\"answer\":\"The thesis targets two interpretability goals: explaining individual model predictions and discovering latent patterns from data.\"},{\"question\":\"How does the thesis explain predictions from black-box sequence models?\",\"answer\":\"It develops a method that approximates a combinatorial objective to clarify the decision-making process of general sequence models.\"},{\"question\":\"What models are developed for political science and labor economics?\",\"answer\":\"For political science, it introduces a text-based ideal point model that quantifies political positions from text. For labor economics, it proposes CAREER transfer learning for career trajectories and uses it to estimate the history-adjusted gender wage gap.\"}]","Interpretable Machine Learning for the Social Sciences - Applications in Political Science and Labor Economics | PDF",1785731868,393,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"interpretable-machine-learning-for-the-social-sciences-applications-in-political-science-and-labor-economics","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/interpretable-machine-learning-for-the-social-sciences-applications-in-political-science-and-labor-economics/120757/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What interpretability goals does the thesis focus on?","Question",{"text":75,"@type":76},"The thesis targets two interpretability goals: explaining individual model predictions and discovering latent patterns from data.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the thesis explain predictions from black-box sequence models?",{"text":80,"@type":76},"It develops a method that approximates a combinatorial objective to clarify the decision-making process of general sequence models.",{"name":82,"@type":73,"acceptedAnswer":83},"What models are developed for political science and labor economics?",{"text":84,"@type":76},"For political science, it introduces a text-based ideal point model that quantifies political positions from text. For labor economics, it proposes CAREER transfer learning for career trajectories and uses it to estimate the history-adjusted gender wage gap.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]