[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-117794-en":3,"doc-seo-117794-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},117794,687197100911,"Himbo","https://ap-avatar.wpscdn.com/avatar/a000239b6f1da00475?x-image-process=image/resize,m_fixed,w_180,h_180&k=1785132997149421697",8,"Research & Report","A Common Misassumption in Online Experiments with Machine Learning Models - Opinion Paper","Online experiments such as Randomised Controlled Trials and A/B-tests are widely used by web platforms to estimate causal effects of switching between system variants on selected metrics. This paper targets the setting where variants correspond to machine learning models and the experiment is treated as the final arbiter for shipping the better model. It argues that key assumptions for unbiased causal effect estimation are rarely satisfied in practice because model interference can arise when variants learn from pooled data, undermining conclusions and motivating implications for both practitioners and the research community.","arXiv :2304 . 10900v1 [ cs .LG] 21 Apr 2023  \nOPINION PAPER  \nA Common Misassumption in Online Experiments with Machine Learning Models  \nOlivier Jeunen  \nShareChat  \nEdinburgh, UK  \n[jeunen@sharechat.co](jeunen@sharechat.co)  \nAbstract  \nOnline experiments such as Randomised Controlled Trials (RCTs) or A/B-tests are the bread and butter of modern platforms on the web. They are conducted continuously to allow platforms to estimate the causal e􀀋ect of replacing system variant \\A\" with variant \\B\", on some metric of interest. These variants can di􀀋er in many aspects. In this paper, we focus on the common use-case where they correspond to machine learning models. The online experiment then serves as the 􀀌nal arbiter to decide which model is superior, and should thus be shipped.  \nThe statistical literature on causal e􀀋ect estimation from RCTs has a substantial history, which contributes deservedly to the level of trust researchers and practitioners have in this \\gold standard\" of evaluation practices. Nevertheless, in the particular case of machine learning experiments, we remark that certain critical issues remain. Speci􀀌cally, the assumptions that are required to ascertain that A/B-tests yield unbiased estimates of the causal e􀀋ect, are seldom met in practical applications. We argue that, because variants typically learn using pooled data, a lack of model interference cannot be guaranteed. This undermines the conclusions we can draw from online experiments with machine learning models. We discuss the implications this has for practitioners, and for the research literature.  \nRandomised Controlled Trials and their Assumptions  \nRandomised experiments have existed in the scienti􀀌c literature for close to 140 years, 􀀌rst introduced in psychology [Peirce and Jastrow, 1884] . Since then, they have been a popular topic of study in the statistical literature |a feat often ascribed to the seminal works of Fisher [1925 , 1936]| and are generally well-understood [Imbens and Rubin, 2015] . Randomised Controlled Trials (RCTs) form the theoretical basis for the online experiments that modern web platforms run continuously [Gupta et al. , 2019], colloquially known as A/B-tests [Kohavi et al. , 2020] .  \nGenerally speaking, RCTs deal with treatments being applied to units, leading to certain outcomes [Rubin, 1974] . Typical examples from the early literature revolve around agricultural applications, where we have types of fertiliser we can apply to plots of land, which has an e􀀋ect on crop yield. In an RCT, we randomly assign units to treatment/control, and as a result, the average measured outcomes for units under control C and treatment T give a 􀀌nite-sample estimate of  \nthe causal e􀀋ect that the treatment has on the outcome, the Average Treatment E􀀋ect (ATE):  \n􀀖 􀀖  \nATE(C ! T ; Y ) = Y(T ) 􀀀 Y(C): (1)  \nUnder seemingly reasonable and light assumptions, this estimate is consistent and unbiased. That is, given in􀀌nite samples, the expectation of the estimate converges to the true average causal e􀀋ect of applying the treatment instead of the control [Imbens and Rubin, 2015] .  \nThe assumption that is needed to support this claim, is called the Stable Unit Treatment Value Assumption (SUTVA) . The SUTVA implies that outcomes of units are independent of the outcomes of other units under di􀀋erent treatments. An agricultural example of when things can go wrong is given by Rubin [1974]: when \\the plots are in such close proximity that following rainfall plot i receives fertilizer [sic] from adjacent plots.\" In many classical cases, it is easily argued that this assumption holds due to experimental design choices and proper randomisation. In A/B-tests, it is often assumed that there are no \\network e􀀋ects\" or \\spillovers\" among units [Gupta et al. , 2019 , §10] . When spillovers are known to exist, alternative experimental design frameworks have been proposed in the literature. These are typically proposed because of some shared resource among tre","cbCaithxZNrhmR5j","https://ap.wps.com/l/cbCaithxZNrhmR5j","pdf",3851564,1,9,"English","en",105,"# Randomised Controlled Trials and their Assumptions\n## Stable Unit Treatment Value Assumption (SUTVA)\n# Online Experiments with Machine Learning Models\n## Contextual bandit setup and A/B allocation\n## Interference via logged data","[{\"question\":\"What common role do online experiments like RCTs and A/B-tests play on web platforms?\",\"answer\":\"They are used continuously to estimate the causal effect of replacing one system variant with another on an outcome metric.\"},{\"question\":\"Why might A/B-test causal estimates be biased when variants are machine learning models?\",\"answer\":\"Because required assumptions for unbiased causal inference are often violated: variants can interfere when they learn using pooled data.\"},{\"question\":\"What is the key assumption SUTVA, and how does this paper relate it to online ML experiments?\",\"answer\":\"SUTVA requires outcomes to be independent across units under different treatments. The paper argues SUTVA can be violated in many everyday ML experiments even without network effects.\"}]","A Common Misassumption in Online Experiments with Machine Learning Models - Opinion Paper | PDF",1785679602,23,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"a-common-misassumption-in-online-experiments-with-machine-learning-models-opinion-paper","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/a-common-misassumption-in-online-experiments-with-machine-learning-models-opinion-paper/117794/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-02",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What common role do online experiments like RCTs and A/B-tests play on web platforms?","Question",{"text":75,"@type":76},"They are used continuously to estimate the causal effect of replacing one system variant with another on an outcome metric.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Why might A/B-test causal estimates be biased when variants are machine learning models?",{"text":80,"@type":76},"Because required assumptions for unbiased causal inference are often violated: variants can interfere when they learn using pooled data.",{"name":82,"@type":73,"acceptedAnswer":83},"What is the key assumption SUTVA, and how does this paper relate it to online ML experiments?",{"text":84,"@type":76},"SUTVA requires outcomes to be independent across units under different treatments. The paper argues SUTVA can be violated in many everyday ML experiments even without network effects.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,127,130,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":21,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]