[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83189-en":3,"doc-seo-83189-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},83189,1374391974468,"Eden","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","Is Randomness Necessary for Adaptive Data Analysis","Adaptive Data Analysis (ADA) studies how to avoid false discovery and overfitting when the same dataset is reused to answer adaptively chosen statistical queries. For a dataset of n i.i.d. samples from an unknown distribution P, the goal is to determine the largest query sequence length k supportable while maintaining ε-accurate answers. The work resolves a longstanding open question on whether randomness is fundamentally required. In the information-theoretic Random Oracle model, deterministic mechanisms fail after k = O(n) queries for unbounded analysts.","arXiv :2607 .07085v 1 [ cs .CR] 8 Jul 2026  \nIs Randomness Necessary for Adaptive Data Analysis?  \nEdith Cohen*,† Haim Kaplan†,* Yishay Mansour†,* Shay Sapir‡,* Uri Stemmer†,*  \nAbstract  \nThe Adaptive Data Analysis (ADA) problem formalizes the challenge of preventing false discovery and overfitting when a dataset is repeatedly reused. Formally, our input is a dataset containing n i.i.d. samples from an unknown distribution P over a domain X , and our goal is to answer a sequence of k adaptively chosen statistical queries with respect to P. The main question is how many queries we can support (i.e., how large k can be), primarily as a function of the number of samples n. This question has been intensively studied and is relatively well-understood for randomized mechanisms: there are computationally efficient mechanisms that support k ≈ n2 queries, and no computationally efficient mechanism can answer k ≫ n2 queries. In this paper, we address a fundamental question: is randomness necessary for ADA?  \nDespite a decade of work on ADA, this question remains open. A folklore observation dating back to the initial works on ADA is that randomness is not necessary when the analyst is computationally bounded. Yet, the necessity of randomness against computationally unbounded analysts has remained elusive. Our main contribution resolves this gap in the information-theoretic Random Oracle model. Perhaps surprisingly, we show that randomness is strictly necessary to answer a non-trivial number of adaptive q˜ueries: when the analyst is unbounded, any deterministic mechanism can be forced to fail after  \njust k = O (n) queries.  \n1 Introduction  \nClassical statistical theory for asserting the validity of a hypothesis tells us that the description of the hypothesis should be independent from the data on which it is evaluated. In practice, however, analysts frequently reuse the same dataset to answer multiple questions. Often, this exploration is adaptive: an analyst asks a statistical question, observes the answer, and uses that information to decide what to ask next. While this is how data analysis naturally happens, it makes it incredibly easy to overfit. Without careful intervention, the analyst will quickly find patterns that exist in the specific sample but fail to generalize to the true underlying distribution.  \nThis problem was formalized by Dwork et al. [DFH+ 15b] in what has come to be known as the Adaptive Data Analysis (ADA) problem. In this problem, a mechanism M holds a dataset S consisting of n samples drawn i.i.d. from an unknown distribution P over a domain X. An analyst iteratively (and adaptively) submits a sequence of k statistical queries q : X → [0 , 1] . The mechanism M must answer these queries such that every answer is ε-accurate (i.e., within ±ε) with respect to the true expectation of the query over P. To provide worst-case guarantees, the analyst is assumed to be adversarial, and is often referred to as an attacker. The central question in ADA is establishing the optimal sample complexity: how large can k beas a function of n?  \nThis question is relatively well-understood for randomized mechanisms. A deep connection between ADA and various stability notions (such as differential privacy) has shown that randomized mechanisms can safely answer a large number of adaptive queries without overfitting [DFH + 15b, DFH+ 15a, BNS+ 16 , RRST16 , RZ16 , FS17 , FS18 , FRR18 , LS19 , SZ20 , JLN+ 20 , SL23 , Bla23] . Specifically, there are computationally efficient randomized mechanisms supporting k ≈ n2 queries, and, without making specific assumptions, no mechanism, even computationally unbounded, can answer k ≫ n2 queries.1  \n* Google Research †Tel Aviv University ‡Weizmann Institute of Science  \n1 If one assumes that the domain X is not too large, namely that n ≥ polylog |X|, then there exist computationally inefficient mechanisms supporting k ≈ exp(n/ polylog |X|) queries. See Section 1.3 (related works) .  \nHowever, all ","cbCaij0EKGGhaSAu","https://ap.wps.com/l/cbCaij0EKGGhaSAu","pdf",714411,1,22,"English","en",105,"# Abstract\n# Introduction\n## Our Contributions","[{\"question\":\"What problem does Adaptive Data Analysis (ADA) formalize?\",\"answer\":\"ADA formalizes the challenge of preventing false discovery and overfitting when a dataset is repeatedly reused to answer adaptively chosen statistical queries.\"},{\"question\":\"How does the paper frame the core question about query complexity?\",\"answer\":\"It asks how large the adaptive query sequence length k can be as a function of the number of samples n while keeping answers ε-accurate with respect to the true distribution P.\"},{\"question\":\"Is randomness necessary for ADA according to the paper?\",\"answer\":\"Yes. In the information-theoretic Random Oracle model, randomness is strictly necessary: for unbounded analysts, any deterministic mechanism can be forced to fail after k = O(n) queries.\"}]",1784185848,55,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"is-randomness-necessary-for-adaptive-data-analysis","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/is-randomness-necessary-for-adaptive-data-analysis/83189/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-20","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does Adaptive Data Analysis (ADA) formalize?","Question",{"text":75,"@type":76},"ADA formalizes the challenge of preventing false discovery and overfitting when a dataset is repeatedly reused to answer adaptively chosen statistical queries.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the paper frame the core question about query complexity?",{"text":80,"@type":76},"It asks how large the adaptive query sequence length k can be as a function of the number of samples n while keeping answers ε-accurate with respect to the true distribution P.",{"name":82,"@type":73,"acceptedAnswer":83},"Is randomness necessary for ADA according to the paper?",{"text":84,"@type":76},"Yes. In the information-theoretic Random Oracle model, randomness is strictly necessary: for unbounded analysts, any deterministic mechanism can be forced to fail after k = O(n) queries.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]