[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-125325-en":3,"doc-seo-125325-105":30,"detail-sidebar-cat-0-en-105":95},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},125325,962075114765,"Quinn","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","Machine learning and the optimization of prediction-based policies - Research article abstract","The document presents an operational method to design and compare prediction-based public policies. It links classification prediction errors to social welfare by defining a loss function as the distance from an ideal error-free policy, producing a ranking measure tied to ROC-curve concepts. The approach extends cost-isometrics to heterogeneous Type I and Type II error costs. Empirical application targets inaccurate Italian self-employed and sole-proprietorship tax returns, showing revenue gains and model insights such as cross-sector heterogeneity and bunching effects.","Technological Forecasting & Social Change 199 (2024) 123080  \n| Machine learning and the optimization of prediction-based policies✩ Pietro Battistona,b,∗, Simona Gamba c, Alessandro Santorod\u003Cbr>a Department of Economics and Management, University of Parma, Italy b Department of Economics and Management, University of Pisa, Italy\u003Cbr>c Department of Economics, Management and Quantitative Methods, Università degli Studi di Milano, Italy d DEMS, University of Milan-Bicocca, Italy |  |  |\n| --- | --- | --- |\n| A R T I C L E I N F O |  | A B S T R A C T |\n| JEL classification:\u003Cbr>C53 D78 H50\u003Cbr>Keywords: Prediction Public policy ROC curve Machine learning Tax behavior |  | We present a procedure for the optimal implementation of public policies that involve predicting an individual behavior or characteristic. By linking prediction errors of any given classification model to the resulting social welfare, we provide a simple measure to rank different models and select the optimal one. Such measure is defined as the difference between the social welfare of a given policy and that of an error-free policy, and it is related to the ROC curve employed in the Machine Learning literature. We extend the cost isometrics approach described in the literature by considering the case of heterogeneous costs of type I and II errors. We apply our approach to the prediction of inaccurate tax returns issued by Italian self-employed and sole proprietorships. We show that the approach can result in substantial increases in revenues, and that random forest models, beyond providing comparatively good predictions, yield important insights. In our case, they both provide empirical support for existing theories on tax evasion—highlighting, for instance, cross-sectoral heterogeneity—and extend our understanding of the phenomenon—such as the role of bunching. |\n\n1. Introduction  \nPolicies that attribute a substantial role to predicting individual behavior or characteristics are acquiring an increasingly important role: there are many policy applications where causal inference isnot central, or even necessary (Kleinberg et al., 2015), and prediction problems bear increasingly large importance for economics and economists (Varian, 2014). Indeed, with prediction-based policies being employed in fields as diverse as schooling systems, police organization, the judicial sector, tax enforcement, and fiscal stimuli, it is clear that their effective deployment deserves a crucial position in the context of public economics.  \nThe goal of the present work is to provide and discuss an operational link between machine learning methods and prediction-based public policies. Indeed, supervised machine learning algorithms, explicitly designed with the goal of optimizing out-of-sample prediction, can be fruitfully employed in implementing policies in the above mentioned fields, as well as in other public services. By taking into account a number of individual features, they can be used to identify as precisely as possible an ideal target, consisting of units on which a given intervention is most cost-effective, in order to target said intervention at them. Hence, we consider a benchmark ideal policy, in which the  \nidentification of the target, and hence the selection of recipients of the intervention, is error-free. Compared to it, the efficiency of any real policy depends on the frequency of both false negatives and false positives. False negatives are units that should be targeted but are not, false positives, vice–versa, are targeted but should not. The estimated loss of social welfare, including implementation costs, depends on the frequency and the relevance of the two types of errors. Hence, we present and characterize a loss function that assigns a social cost to each prediction-based policy, such cost being the distance from the ideal policy. We show how to compare different methods for the selection of the target sample; our approach can employ any kind of predictive m","cbCaiok5FVKT3nNV","https://ap.wps.com/l/cbCaiok5FVKT3nNV","pdf",1120030,1,16,"English","en",105,"# Introduction\n## Prediction-based public policy and motivation\n## Operational link between machine learning and policy design\n## Loss function based on social welfare and classification errors\n## Empirical application to tax return prediction","[{\"question\":\"What is the paper’s main goal regarding prediction-based policies?\",\"answer\":\"To provide a practical link between supervised machine learning and prediction-based public policies by optimizing which recipients to target based on prediction risk scores.\"},{\"question\":\"How does the proposed method evaluate and compare different predictive models?\",\"answer\":\"It defines a social-cost loss measuring the distance from an ideal error-free policy, so models can be ranked by how much their associated policy welfare differs from the error-free benchmark.\"},{\"question\":\"How is the method extended beyond standard cost-sensitive approaches?\",\"answer\":\"It extends the cost isometrics framework by allowing heterogeneous costs for Type I and Type II classification errors, then relates these costs to policy social welfare.\"},{\"question\":\"What does the empirical application to tax returns show?\",\"answer\":\"Applied to predicting inaccurate tax returns, the approach can substantially increase revenues and, using models like random forests, supports theories on tax evasion while highlighting cross-sector heterogeneity and the role of bunching.\"}]","Machine learning and the optimization of prediction-based policies - Research article abstract | PDF",1785898189,40,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":90,"head_meta":92,"extra_data":94,"updated_unix":28},"machine-learning-and-the-optimization-of-prediction-based-policies-research-article-abstract","",{"@graph":36,"@context":89},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/machine-learning-and-the-optimization-of-prediction-based-policies-research-article-abstract/125325/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-05",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81,85],{"name":72,"@type":73,"acceptedAnswer":74},"What is the paper’s main goal regarding prediction-based policies?","Question",{"text":75,"@type":76},"To provide a practical link between supervised machine learning and prediction-based public policies by optimizing which recipients to target based on prediction risk scores.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the proposed method evaluate and compare different predictive models?",{"text":80,"@type":76},"It defines a social-cost loss measuring the distance from an ideal error-free policy, so models can be ranked by how much their associated policy welfare differs from the error-free benchmark.",{"name":82,"@type":73,"acceptedAnswer":83},"How is the method extended beyond standard cost-sensitive approaches?",{"text":84,"@type":76},"It extends the cost isometrics framework by allowing heterogeneous costs for Type I and Type II classification errors, then relates these costs to policy social welfare.",{"name":86,"@type":73,"acceptedAnswer":87},"What does the empirical application to tax returns show?",{"text":88,"@type":76},"Applied to predicting inaccurate tax returns, the approach can substantially increase revenues and, using models like random forests, supports theories on tax evasion while highlighting cross-sector heterogeneity and the role of bunching.","https://schema.org",{"og:url":52,"og:type":91,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":93,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":96},[97,101,105,109,114,119,123,126,131,134,138],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":106,"show_sort_weight":107,"slug":108},"Exam",70,"exam",{"id":110,"doc_module":4,"doc_module_name":46,"category_name":111,"show_sort_weight":112,"slug":113},5,"Comic",60,"comic",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},6,"Technology",50,"technology",{"id":120,"doc_module":4,"doc_module_name":46,"category_name":121,"show_sort_weight":29,"slug":122},7,"Healthcare","healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":124,"slug":125},30,"research-report",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":129,"slug":130},9,"Religion & Spirituality",20,"religion-spirituality",{"id":129,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":129,"slug":133},"World Cup","world-cup",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":135,"slug":137},10,"Lifestyle","lifestyle",{"id":139,"doc_module":4,"doc_module_name":46,"category_name":140,"show_sort_weight":110,"slug":141},19,"General","general"]