[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-123012-en":3,"doc-seo-123012-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},123012,8796095462418,"Noah","https://ap-avatar.wpscdn.com/avatar/80000253c1241d02b47?x-image-process=image/resize,m_fixed,w_180,h_180&k=1778826106357471780",8,"Research & Report","Robust Optimization for Inference on Machine Learning Generated Variables","Supervised machine learning (SML) is increasingly used to convert unstructured text and images into measurable variables for regression-based inference and theory testing. Because SML-derived variables contain measurement errors relative to the underlying constructs, these errors can bias coefficient estimates and distort hypothesis tests. This study develops a robust optimization approach for a generalized regression setting with heteroscedastic, unknown measurement errors, integrating an uncertainty set learned from labeled data. The method proves consistency and efficiency, and simulation results confirm reduced bias and improved inference accuracy.","Proceedings of the 57th Hawaii International Conference on System Sciences | 2024  \nRobust Optimization for Inference on Machine Learning Generated  \nVariables  \nAaron Schecter University of Georgia  [aschecter@uga.edu](aschecter@uga.edu)  \nWeifeng Li University of Georgia  [weifeng.li@uga.edu](weifeng.li@uga.edu)  \nAbstract  \nLeveraging supervised machine learning (SML) algorithms to operationalize constructs from unstructured data like text or images is becoming common in practice and research. As a result, variables generated through SML are used in regression models to make inferences and test theories. However, variables produced by SML will have measurement errors compared to the underlying construct. We propose using robust optimization to reduce the negative impact of these errors and produce less biased coefficient estimates while conducting more accurate hypothesis testing. To extend the burgeoning literature on this issue, our proposed method focuses on the generalized research setting where a flexible number of dependent and independent variables are measured by SML algorithms. We combine recent robust optimization techniques tofit a linear regression model in the presence of uncertain measurement error. We theoretically demonstrate the consistency and efficiency of the robust approach. Through simulations, we demonstrate the effectiveness of our approach.  \nKeywords: Robust optimization, machine learning, statistical inference, regression.  \n1. Introduction  \nWith the increasing availability of unstructured data such as text and images, the information systems research community has grown a significant interest in operationalizing constructs from unstructured data using supervised machine learning (SML) methods. SML methods estimate measures through learning the  \nunderlying relationships between constructs of interests (e.g., sentiments) and unstructured data (e.g., customer reviews) . Applications of SML methods include determining the quality of images (Zhang, Lee, Singh,& Srinivasan, 2022), classifying the sentiment of product reviews (Tirunillai & Tellis, 2012), predicting post quality in stock discussion boards (Gu, Konana, Rajagopalan, & Chen, 2007), and more. SML methods have shown great potential for providing reliable measurements for constructs of theoretical or practical importance.  \nTo incorporate SML-based variables into econometric analysis, past hybrid studies often use a two-step estimation framework (Qiao & Huang, 2021): in the first step, SML methods are used to develop measures of interest from unstructured data such as text and images; in the second step, these SML-based measures are included in an empirical regression model. Such a two-step estimation framework expands the scope of the information systems research community by allowing researchers to examine phenomena and test theories in previously unquantifiable contexts. In the two-step estimation framework, the variables generated by SML generally have measurement errors, originating from SML methods’ imperfect estimates of the target constructs (Yang, Adomavicius, Burtch, & Ren, 2018; Qiao & Huang, 2021) . For example, when using text mining SML to predict customer satisfaction from textual reviews, a researcher might face the risk of mistaking a satisfied customer for an unsatisfied one or vice versa. Such measurement errors can cause biases in the second step estimation. Specifically, measurement errors in the first step can attenuate or amplify the coefficient estimates of SML-based variables and further distort the estimation of the dependent variable in the second step (Carroll, Ruppert, Stefanski, &  \nURI: [https://hdl.handle.net/10125/106509](https://hdl.handle.net/10125/106509)[ ](https://hdl.handle.net/10125/106509)[978-0-9981331-7-1](978-0-9981331-7-1)  \n(CC BY-NC-ND 4 .0)  \nPage 1100  \nCrainiceanu, 2006) . In response to measurement errors, several error correction techniques have been proposed, including the method of moment","cbCaikJBfcKmDE5G","https://ap.wps.com/l/cbCaikJBfcKmDE5G","pdf",353017,1,10,"English","en",105,"# 1. Introduction\n## Motivation and two-step estimation with SML-generated variables\n## Measurement error challenges\n## Proposed robust optimization approach","[{\"question\":\"Why do SML-generated variables cause problems in econometric regression inference?\",\"answer\":\"SML variables have measurement errors because the learned measures imperfectly estimate the underlying constructs. These errors can attenuate or amplify coefficient estimates and distort the dependent-variable estimation in the second-step regression.\"},{\"question\":\"What robust optimization method is proposed to address measurement errors?\",\"answer\":\"The method formulates a robust optimization model that accounts for heteroscedastic, unknown measurement errors as adversarial effects. It incorporates an uncertainty set learned from labeled data directly into the regression-fitting process.\"},{\"question\":\"How is the proposed approach evaluated and what are the results?\",\"answer\":\"The study derives theoretical guarantees showing consistency and efficiency of the robust method, and then runs simulation experiments. Simulations show that using the robust method with a correction term yields less biased coefficient estimates than OLS.\"}]","Robust Optimization for Inference on Machine Learning Generated Variables | PDF",1785814175,25,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"robust-optimization-for-inference-on-machine-learning-generated-variables","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/robust-optimization-for-inference-on-machine-learning-generated-variables/123012/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why do SML-generated variables cause problems in econometric regression inference?","Question",{"text":75,"@type":76},"SML variables have measurement errors because the learned measures imperfectly estimate the underlying constructs. These errors can attenuate or amplify coefficient estimates and distort the dependent-variable estimation in the second-step regression.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What robust optimization method is proposed to address measurement errors?",{"text":80,"@type":76},"The method formulates a robust optimization model that accounts for heteroscedastic, unknown measurement errors as adversarial effects. It incorporates an uncertainty set learned from labeled data directly into the regression-fitting process.",{"name":82,"@type":73,"acceptedAnswer":83},"How is the proposed approach evaluated and what are the results?",{"text":84,"@type":76},"The study derives theoretical guarantees showing consistency and efficiency of the robust method, and then runs simulation experiments. Simulations show that using the robust method with a correction term yields less biased coefficient estimates than OLS.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":21,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":21,"slug":133},"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]