[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-116892-en":3,"doc-seo-116892-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},116892,687197100911,"Himbo","https://ap-avatar.wpscdn.com/avatar/a000239b6f1da00475?x-image-process=image/resize,m_fixed,w_180,h_180&k=1785132997149421697",8,"Research & Report","Measuring Human Capital with Social Media Data and Machine Learning","The paper addresses persistent gaps in survey data for tracking socio-economic development by using geolocated Twitter information and machine learning to estimate human capital outcomes. It constructs interpretable indicators at municipality and county levels from 25 million+ tweets, combining penetration, usage, text quality, topics, sentiment, and network measures, with cluster-neighborhood estimates for feature enrichment and imputation. Stacking models tuned via grid search and assessed with five-fold cross-validation predict educational attainment with about 70% explained variation in Mexico and 65% in the US.","Faculty of Business, Economics and Social Sciences  \nDepartment of Social Sciences  \nUniversity of Bern Social Sciences Working Paper No. 46  \nMeasuring Human Capital with Social Media Data and Machine Learning  \nsource: [https://doi.org/10.48350/182366 | downloaded:](https://doi.org/10.48350/182366 | downloaded:) 15.5.2023  \nMartina Jakob and Sebastian Heinrich  \nMay 5 , 2023  \n[http://ideas. repec.org/p/bss/wpaper/46.html](http://ideas. repec.org/p/bss/wpaper/46.html)  \n[http://econpapers. repec.org/paper/bsswpaper/46.htm](http://econpapers. repec.org/paper/bsswpaper/46.htm)  \nUniversity of Bern Department of Social Sciences Fabrikstrasse 8  \nCH-3012 Bern  \nTel. +41 (0)31 684 48 11 Fax +41 (0)31 684 48 17 [info@sowi. unibe.ch](info@sowi. unibe.ch)[ ](info@sowi. unibe.ch)[www.sowi. unibe.ch](www.sowi. unibe.ch)  \nMeasuring Human Capital with Social Media Data and  \nMachine Learning  \nMartina Jakob  \nUniversity of Bern  \n[martina.jakob@unibe.ch](martina.jakob@unibe.ch)  \nSebastian Heinrich  \nETH Zurich  \nheinrich@kof.ethz.ch  \nMay 5, 2023  \nIn response to persistent gaps in the availability of survey data, a new strand of research leverages alternative data sources through machine learning to track global development. While previous applications have been successful at predicting outcomes such as wealth, poverty or population density, we show that educational outcomes can be accurately estimated using geo-coded Twitter data and machine learning. Based on various input features, including user and tweet characteristics, topics, spelling mistakes, and network indicators, we can account for ∼70 percent of the variation in educational attainment in Mexican municipalities and US counties.  \nKeywords: machine learning, social media data, education, human capital, indicators, natural language processing  \nJEL Codes: C53, C80, O11, O15, I21, I25  \nWe are grateful to Ben Jann, Carla Coccia, Mauricio Romero, and Joel Ferguson for their helpful comments and suggestions.  \n1 Introduction  \nReliable data on key socio-economic outcomes enables policy-makers to take informed decisions and promote societal development. However, many countries are plagued by a pervasive lack of such data, limiting their ability to track progress and evaluate policies. To address the problem, a growing strand of literature uses alternative data sources such as satellite imagery or phone records to bridge the existing gaps in data availability (Burke et al., 2021) . While previous studies have successfully predicted outcomes such as wealth, income or population density, this paper proposes an innovative approach to measuring human capital using geolocated Twitter data.  \nSpecifically, we construct a series of interpretable measures of human capital at low administrative units (municipality in Mexico and county in the United States) based on over 25 million tweets. Our feature matrix includes simple Twitter penetration (e.g., user densities) and usage statistics (e.g., tweet length), text-based indicators on spelling mistakes (e.g. , frequency of grammar mistakes), topics, (e.g., share of tweets about science) and sentiments (e.g., share of negative tweets) as well as network indicators (e.g., closeness centrality) . For each input, we compute cluster-level estimates based on geographical neighbors, and use them both as additional features and to impute missing values. We then train a stacking regressor combining five machine learning algorithms—elastic net regression, gradient boosting, support vector regression, nearest neighbor regression, and a feed-forward neural network — to predict educational attainment for Mexican municipalities (N = 2,457) and US counties (N = 3,141) . We apply grid search to tune the relevant hyperparameters of each model, and evaluate the performance of the final models using five-fold cross-validation.  \nOur predictions account for 70 percent of the variation in years of schooling in Mexican municipalities and 65 percent in US counti","cbCaihfWJtsOw2Nt","https://ap.wps.com/l/cbCaihfWJtsOw2Nt","pdf",4661938,1,39,"English","en",105,"# Introduction\n## Data gap and alternative data sources\n## Research approach and feature construction\n## Modeling strategy and evaluation\n## Key results and predictors\n## Challenges and robustness checks","[{\"question\":\"How does the study measure human capital using social media data?\",\"answer\":\"It builds interpretable human-capital measures from geolocated Twitter text and metadata at municipality (Mexico) and county (US) levels, then derives cluster-level neighborhood estimates for feature enrichment and imputation.\"},{\"question\":\"Which machine learning models are used to predict educational outcomes?\",\"answer\":\"The paper trains a stacking regressor that combines five algorithms: elastic net regression, gradient boosting, support vector regression, nearest neighbor regression, and a feed-forward neural network.\"},{\"question\":\"What level of predictive accuracy is achieved for educational attainment?\",\"answer\":\"Predictions explain about 70% of the variation in years of schooling for Mexican municipalities and about 65% for US counties, with strong results especially for higher education levels.\"}]","Measuring Human Capital with Social Media Data and Machine Learning | PDF",1785672295,98,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"measuring-human-capital-with-social-media-data-and-machine-learning","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/measuring-human-capital-with-social-media-data-and-machine-learning/116892/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-02",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"How does the study measure human capital using social media data?","Question",{"text":75,"@type":76},"It builds interpretable human-capital measures from geolocated Twitter text and metadata at municipality (Mexico) and county (US) levels, then derives cluster-level neighborhood estimates for feature enrichment and imputation.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Which machine learning models are used to predict educational outcomes?",{"text":80,"@type":76},"The paper trains a stacking regressor that combines five algorithms: elastic net regression, gradient boosting, support vector regression, nearest neighbor regression, and a feed-forward neural network.",{"name":82,"@type":73,"acceptedAnswer":83},"What level of predictive accuracy is achieved for educational attainment?",{"text":84,"@type":76},"Predictions explain about 70% of the variation in years of schooling for Mexican municipalities and about 65% for US counties, with strong results especially for higher education levels.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]