[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-134458-en":3,"doc-seo-134458-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},134458,1099523885336,"Taylor Morgan","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","Conformal Frequency Estimation with Sketched Data - NeurIPS 2022","A flexible conformal inference framework constructs frequentist confidence intervals for object frequencies in very large data sets using only a much smaller sketch. The method is data-adaptive and does not require knowledge of the underlying data distribution or the details of the sketching procedure. It guarantees provable validity under the single assumption of data exchangeability. The paper emphasizes applications with the count-min sketch and a nonlinear variation, comparing results against frequentist and Bayesian baselines via simulations and experiments on SARS-CoV-2 DNA sequences and English literature.","Conformal Frequency Estimation with Sketched Data  \nMatteo Sesia  \nDepartment of Data Sciences and Operations University of Southern California Los Angeles, California, USA [sesia@marshall.usc.edu](sesia@marshall.usc.edu)  \nStefano Favaro  \nDepartment of Economics and Statistics University of Torino and Collegio Carlo Alberto Torino, Italy stefano .favaro@unito .it  \nAbstract  \nA ﬂexible conformal inference method is developed to construct conﬁdence intervals for the frequencies of queried objects in very large data sets, based on amuch smaller sketch of those data. The approach is data-adaptive and requires no knowledge of the data distribution or of the details of the sketching algorithm;  \ninstead, it constructs provably valid frequentist conﬁdence intervals under the sole assumption of data exchangeability. Although our solution is broadly applicable, this paper focuses on applications involving the count-min sketch algorithm anda non-linear variation thereof. The performance is compared to that of frequentistand Bayesian alternatives through simulations and experiments with data sets of SARS-CoV-2 DNA sequences and classic English literature.  \n1 Introduction  \n1.1 Frequency queries from sketched data  \nAn important task in computer science is to estimate the frequency of an object given a lossy compressed representation, or sketch, of a big data set [1, 2]; this task has real-world relevance in diverse ﬁelds including machine learning [3], cybersecurity [4], natural language processing [5], genetics [6], and privacy [7] . Practically, sketching may be motivated for example by memory limitations, as large numbers of distinct symbols may otherwise be computationally expensive to analyze [6], or by privacy constraints, in situations where the original data contain sensitive information [8] . There exist many sketching algorithms, several of which are speciﬁcally designed to enable efﬁcient approximations of the empirical frequencies of the compressed objects. We refer to the monograph of [9] for a recent review of sketching. Building upon this literature, our paper studies the problem of precisely quantifying the uncertainty of empirical frequency estimates obtained from sketched data without making strong assumptions about the inner workings ofthe sketching procedure, which maybe complex and even unknown. The key ideas of the described solution are in principle applicable regardless of how the data are compressed, but the exposition of this paper will focus for simplicity on a particularly well-known sketching algorithm and some widely applied variations thereof.  \n1.2 The count-min sketch  \nThe count-min sketch (CMS) of [10] is a renowned algorithm for compressing a data set of m objects Z1 ; : : : ; Zm 2 Z , with Z being a discrete (and possibly inﬁnite) space, into a representation with reduced memory footprint, while allowing simple approximate queries about the observed frequency of any possible z 2 Z. At the heart of the CMS lie d 􀀕 1 different w-wide hash functions hj : Z ! [w] := f1; : : : ; wg, for all j 2 [d] := f1; : : : ; dg and some integer w 􀀕 1. Each hash function maps the elements of Z into one of w possible buckets, and it is designed to ensure that distinct values of z populate the buckets uniformly. Hash functions are typically chosen at random from a pairwise independent family H, which ensures the probability (over the randomness in the  \n36th Conference on Neural Information Processing Systems (NeurIPS 2022) .  \nchoice of hash functions) that two distinct objects z1 ; z2 2 Z are mapped by two different hash functions into the same bucket is equal to 1=w2. The data set Z1 ; : : : ; Zm is then compressed into a sketch matrix C 2 Nd􀀂w with row sums equal to m. The element in the j-th row and k-th column of C counts the number of data points mapped by the j-th hash function into the k-th bucket:  \nm  \nCj;k = X 1 [hj (Zi ) = k] ; j 2 [d]: (1)  \ni=1  \nBecause d and w are such that d 􀀁 w 􀀜 m, the matrix C lo","cbCaifcroTDDqfjw","https://ap.wps.com/l/cbCaifcroTDDqfjw","pdf",261095,1,14,"English","en",105,"# Introduction\n## Frequency queries from sketched data\n## The count-min sketch\n## Uncertainty estimation for sketching with random data","[{\"question\":\"What problem does the paper address?\",\"answer\":\"It builds valid confidence intervals for the true empirical frequency of queried objects when only a lossy sketch of the data is available.\"},{\"question\":\"What assumptions are required for validity?\",\"answer\":\"The approach guarantees frequentist coverage under the sole assumption that the data are exchangeable.\"},{\"question\":\"Which sketching methods are重点研究?\",\"answer\":\"The paper focuses on the count-min sketch algorithm and a non-linear variation of it, and evaluates performance against frequentist and Bayesian alternatives.\"}]","Conformal Frequency Estimation with Sketched Data - NeurIPS 2022 | PDF",1787262791,35,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"conformal-frequency-estimation-with-sketched-data-neurips-2022","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/conformal-frequency-estimation-with-sketched-data-neurips-2022/134458/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-23","2026-08-20",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What problem does the paper address?","Question",{"text":76,"@type":77},"It builds valid confidence intervals for the true empirical frequency of queried objects when only a lossy sketch of the data is available.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"What assumptions are required for validity?",{"text":81,"@type":77},"The approach guarantees frequentist coverage under the sole assumption that the data are exchangeable.",{"name":83,"@type":74,"acceptedAnswer":84},"Which sketching methods are重点研究?",{"text":85,"@type":77},"The paper focuses on the count-min sketch algorithm and a non-linear variation of it, and evaluates performance against frequentist and Bayesian alternatives.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":46,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":46,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]