[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84667-en":3,"doc-seo-84667-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},84667,4810365810221,"Aurora","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","Selectivity Estimation for Linear Queries via Online Learning","Selectivity estimation is a core component of database query optimization, yet learning-based estimators have mainly been analyzed under fixed query distributions on static data. This work develops an online learning framework that allows both the database and query workload to change arbitrarily over time. Using regret as the performance metric, the study derives upper and lower regret bounds for histogram-based linear queries under standard loss functions in both static and dynamic settings.","arXiv :2607 .02895v 1 [ cs .DB] 3 Jul 2026  \nSelectivity Estimation for Linear Queries via Online Learning  \nFANGZHU SHEN, Duke University, USA DEBMALYA PANIGRAHI, Duke University, USA SUDEEPA ROY, Duke University, USA  \nLearning-based approaches for selectivity estimation in databases have gained significant traction in recent years. However, theoretical studies of these learning-based approaches are essentially limited to fixed query distributions on static databases. In practice, both the underlying database and the query workload can dynamically change over time. In this work, we propose an algorithmic framework for learning selectivity of queries in this more general dynamic setup. Inspired by online learning, we measure the performance of the learning algorithm in this setting by its regret, which compares the cumulative loss incurred by the learning algorithm to that of the best fixed strategy. We establish upper and lower bounds on regret for histogram-based linear queries, such as point, range, and subset selection queries, under standard loss functions, in both static and dynamic database settings.  \nAdditional Key Words and Phrases: Selectivity estimation; Online learning; Regret analysis; Loss function; Dynamic data; Linear queries  \n1 Introduction  \nSelectivity estimation is a core component of query optimization, enabling the query optimizer to choose efficient execution plans [19, 20]. Traditional selectivity estimators often make strong assumptions such as independence among attributes on precomputed statistics (e.g., histograms, sketches, and samples), which may not hold in practice [8, 26, 29, 32, 33]. As a powerful alternative, a large number of learning-based approaches to selectivity estimation have emerged in recent years and have exhibited good empirical performance [11, 18, 22, 27, 28, 30, 31, 36, 43]. By leveraging machine learning models to capture complex data distributions and query patterns, these methods have potential for significantly improving accuracy compared to traditional methods.  \nDespite these strong empirical results, theoretical understanding of learning-based approaches for selectivity estimation remains limited. Prior theoretical studies, such as the work by Hu et al. [18], established that selectivity functions for range queries are learnable under the agnostic learning model [13], a generalization of the classical Probably Approximately Correct (PAC) framework [38] . These results guarantee that, given enough training samples, a model’s expected error will be small, if both training and future queries are drawn from the same fixed probability distribution. Recent works have extended these guarantees to handle out-of-distribution queries and sequential data insertions [42, 45, 46] .  \nThese existing results are limited to static or mildly evolving datasets, and heavily rely on stochastic assumptions. Real-world database environments, however, are frequently dynamic: data is continuously inserted, updated, and deleted, and query patterns may shift unpredictably without following any stationary distribution. Under such dynamics, static models trained on historical data can become stale, leading to degraded accuracy or necessitating expensive retraining [39] . While practical systems attempt to handle these dynamics through periodic retraining or lightweight adaptation [24, 30], a rigorous theoretical framework that accommodates these fully dynamic, distribution-free settings would enable principled solutions.  \nIn this work, we present a theoretical study of selectivity estimation in the online learning framework [6, 14, 34]. This framework allows both data and incoming queries to change arbitrarily over time, and does not impose any distributional assumptions on either. The online learning framework naturally mirrors the sequential nature of query processing, where the learner operates in sequential rounds and the prediction must be made before the true selectivity is revealed.  \n2 F","cbCaisv3RjFRdEaW","https://ap.wps.com/l/cbCaisv3RjFRdEaW","pdf",879569,1,29,"English","en",105,"# Introduction\n## Database and Query Models","[{\"question\":\"What limitation of existing learning-based selectivity estimation is addressed in this work?\",\"answer\":\"Prior theory largely assumes fixed query distributions on static databases, while real systems experience fully dynamic data and shifting query patterns.\"},{\"question\":\"How does the paper evaluate the proposed online learning approach?\",\"answer\":\"It measures performance using regret, comparing cumulative loss of the learner to the cumulative loss of the best fixed strategy in hindsight.\"},{\"question\":\"What types of queries are analyzed for regret bounds?\",\"answer\":\"The paper focuses on histogram-based linear queries such as point, range, and subset selection queries, deriving regret bounds in both static and dynamic database settings.\"}]",1784197568,73,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"selectivity-estimation-for-linear-queries-via-online-learning","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/selectivity-estimation-for-linear-queries-via-online-learning/84667/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What limitation of existing learning-based selectivity estimation is addressed in this work?","Question",{"text":75,"@type":76},"Prior theory largely assumes fixed query distributions on static databases, while real systems experience fully dynamic data and shifting query patterns.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the paper evaluate the proposed online learning approach?",{"text":80,"@type":76},"It measures performance using regret, comparing cumulative loss of the learner to the cumulative loss of the best fixed strategy in hindsight.",{"name":82,"@type":73,"acceptedAnswer":83},"What types of queries are analyzed for regret bounds?",{"text":84,"@type":76},"The paper focuses on histogram-based linear queries such as point, range, and subset selection queries, deriving regret bounds in both static and dynamic database settings.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]