[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-124846-en":3,"doc-seo-124846-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},124846,4398048949847,"Eliana","https://ap-avatar.wpscdn.com/avatar/400002536579ef2da7f?_k=1778318612642679267",8,"Research & Report","Learning Decision Policies with Instrumental Variables through Double Machine Learning","Learning decision policies from data-rich offline datasets often suffers from spurious correlations created by hidden confounders. Instrumental variable (IV) regression uses an unconfounded instrument to learn causal relationships among actions, outcomes, and context. Recent IV methods plug a first-stage DNN estimator into a second-stage DNN, which can incur substantial bias, especially under regularisation bias. DML-IV reduces this two-stage bias via a double/debiased machine learning objective, offering fast convergence and O(N−1/2) suboptimality, and outperforming state-of-the-art IV regression on benchmarks.","Learning Decision Policies with Instrumental Variables through Double Machine Learning  \nDaqian Shao 1 , Ashkan Soleymani2 , Francesco Quinzan 1 , Marta Kwiatkowska 1  \n1 Department of Computer Science, University of Oxford  \n2 Department of Electrical Engineering and Computer Science, Massachusetts Institute of Technology  \nAbstract  \nA common issue in learning decision-making policies in data-rich settings is spurious correlations in the offline dataset, which can be caused by hidden confounders. Instrumental variable (IV) regression, which utilises a key unconfounded variable known as the instrument, is a standard technique for learning causal relationships between confounded action, outcome, and context variables. Most recent IV regression algorithms use a two-stage approach, where a deep neural network (DNN) estimator learnt in the first stage is directly plugged into the second stage, in which another DNN is used to estimate the causal effect. Naively plugging the estimator can cause heavy bias in the second stage, especially when regularisation bias is present in the first stage estimator. We propose DML-IV, a non-linear IV regression method that reduces the bias in two-stage IV regressions and effectively learns high-performing policies. We derive a novel learning objective to reduce bias and design the DML-IV algorithm following the double/debiased machine learning (DML) framework. The learnt DML-IV estimator has strong convergence rate and O (N−1/2) suboptimality guarantees that match those when the dataset is unconfounded. DML-IV outperforms state-of-the-art IV regression methods on IV regression benchmarks and learns high-performing policies in the presence of instruments.  \n1 Introduction  \nRecent advances in deep learning (DL) have greatly facilitated the learning of decision-making policies in data-rich settings, but they often lack optimality guarantees. A common issue for learning from offline observational data is the existence of spurious correlations, which are relationships between variables that appear to be causal, but in fact are not. For example, suppose we have aeroplane ticket sales and pricing data in a ticket demand scenario Hartford et al. [2017], and we wish to learn a policy from this offline data that maximises revenue. During holiday season, observational data may contain evidence of a concurrent surge in both ticket sales and prices, which may result in the learning algorithm to learn an incorrect policy that higher ticket prices will drive higher sales.  \nSpurious correlations are often caused by hidden confounders Pearl [2000], which are unobserved variables that influence both the actions (or interventions) and the outcome. In the aeroplane ticket example, the occurrence of popular events and holidays serves as a hidden confounder that raises both ticket prices (actions) and sales (outcome) . To properly account for these hidden  \nconfounders and understand the true causal effect of actions, we need to model the causal (or structural) relationship between the action and the outcome, which is expressed through a causal function. However, learning the causal function in the presence of hidden confounders is known to be challenging and sometimes infeasible Shpitser and Pearl [2008] .  \nA popular approach to deal with hidden confounders is via instrumental variables (IVs) Wright [1928], which are heterogeneous random variables that only affect the action, but not the outcome. These IVs have been used extensively to identify the causal effect of actions in many applications, including econometrics Angrist and Pischke [2009], Reiersöl [1945], drug testings Angrist et al. [1996], and social sciences Angrist [1990] . In the aeroplane ticket example, we can employ supply cost-shifters (e.g., fuel price) as instrumental variables, as their variations are independent of the demand for aeroplane tickets and affect sales solely via ticket prices Blundell et al. [2012] .  \nWe focus on the problem of lea","cbCailXdGzQOGPhv","https://ap.wps.com/l/cbCailXdGzQOGPhv","pdf",1184430,1,35,"English","en",105,"# Introduction\n## Spurious correlations and hidden confounders\n## Instrumental variables for causal identification\n## Offline IV bandit formulation\n## Classical and non-linear IV regression\n## Double/Debiased Machine Learning (DML) framework","[{\"question\":\"What problem does IV regression address in learning decision policies from offline data?\",\"answer\":\"It aims to recover causal effects despite spurious correlations caused by hidden confounders in observational datasets.\"},{\"question\":\"Why can naive two-stage IV regression with neural networks be biased?\",\"answer\":\"Bias can enter the second stage when a first-stage DNN estimator is plugged in, particularly under regularisation bias, slowing convergence of the causal-function estimator.\"},{\"question\":\"How does DML-IV improve over existing two-stage IV regression methods?\",\"answer\":\"DML-IV uses a DML-based learning objective with a Neyman orthogonal score and a cross-fitting regime to reduce two-stage bias and provide convergence guarantees.\"}]","Learning Decision Policies with Instrumental Variables through Double Machine Learning | PDF",1785894973,88,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"learning-decision-policies-with-instrumental-variables-through-double-machine-learning","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/learning-decision-policies-with-instrumental-variables-through-double-machine-learning/124846/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-05",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does IV regression address in learning decision policies from offline data?","Question",{"text":75,"@type":76},"It aims to recover causal effects despite spurious correlations caused by hidden confounders in observational datasets.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Why can naive two-stage IV regression with neural networks be biased?",{"text":80,"@type":76},"Bias can enter the second stage when a first-stage DNN estimator is plugged in, particularly under regularisation bias, slowing convergence of the causal-function estimator.",{"name":82,"@type":73,"acceptedAnswer":83},"How does DML-IV improve over existing two-stage IV regression methods?",{"text":84,"@type":76},"DML-IV uses a DML-based learning objective with a Neyman orthogonal score and a cross-fitting regime to reduce two-stage bias and provide convergence guarantees.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]