[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85268-en":3,"doc-seo-85268-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85268,1374391974585,"Genevieve","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","Efficient Online Proportional Sampling with Applications to Smoothed Online Learning","Efficient online proportional sampling is studied for a high-dimensional domain against a σ-smoothed adaptive adversary. Sampling comes from a dynamically evolving weight function induced by a growing sequence of piecewise-structured partitions, modeled via updates on regions separated by axis-parallel hyperplanes. A new data structure supports efficient updates and proportional sampling without explicit enumeration of the O(td) induced subregions. Tight O(√σT) depth bounds hold under σ-smoothed adversaries, with O(log T) under random-order adversaries. The framework yields no-regret online learning with piecewise-structured rewards under full-information and bandit feedback, with sublinear regret.","arXiv :2607 . 10963v 1 [ cs .LG] 13 Jul 2026  \nEfficient Online Proportional Sampling with Applications to  \nSmoothed Online Learning  \nAmirmahdi Mirfakhar 1 , Maria-Florina Balcan2 , and Hedyeh Beyhaghi 1  \n1 University of Massachusetts Amherst,  \n2 Carnegie Mellon University,  \nAbstract  \nWe study the problem of efficient online proportional sampling from a high-dimensional domain under a σ-smoothed adversary, where the sampling distribution is induced by a dynamically evolving weight function defined over a sequence of piecewise-structured partitions. This setting captures a broad range of applications, including principal-agent games (e.g., pricing and contract design), and algorithm configuration and parameter tuning. The central challenge is maintaining an efficient data structure as the induced partition grows increasingly complex over time—naively, the number of subregions can grow as O (td ) by round t in d dimensions. We design a data structure that supports efficient updatesand proportional sampling while avoiding the cost of explicitly maintaining this exponential growth, where the discontinuities are structured from axis-parallel hyperplanes. Under a σ-smoothed adaptive adversary, we prove a tight O ( √σT) bound on the depth of our data structure, and an O(log T) bound under a random-order adversary—to our knowledge, the first such results for this class of problems. We apply this framework to online learning with piecewise-structured rewards, obtaining efficient no-regret algorithms under both full-information and bandit feedback, with provable sublinear regret guarantees.  \n1 Introduction  \nMany sequential decision-making problems require a learner to repeatedly select actions from a continuous domain in response to an environment whose feedback is evolving over time. A natural and recurring structure in such settings is piecewise-continuous feedback: the reward function is piecewise-Lipschitz within regions of the action space, but changes abruptly across region boundaries. This structure arises organically in a broad range of applications—in principal-agent settings such as dynamic pricing [Balcan and Beyhaghi, 2024, Blum and Hartline, 2005] and contract design [Zhu et al., 2023], where the agent’s best response to the learner’s action creates sharp discontinuities in the reward function, and in algorithm configuration and parameter tuning [Balcan et al. , 2018, 2022, Balcan and Sharma, 2021, Gupta and Roughgarden, 2016], where an algorithm’s overall performance over instances as a function of its parameters is naturally piecewisecontinuous.  \nUnderlying all of these settings is a common computational primitive: proportional sampling from a dynamically evolving weight function. At each round, the learner must select an action with probability proportional to the cumulative rewards observed so far — a distribution that changes as new feedback arrives and the partition of the domain is refined. Proportional sampling is well-studied in static settings [Cochran, 1977, Brewer and Hanif, 1983, Cheung, 2014], and plays a central role in randomized algorithms [Motwani and Raghavan, 1995], online learning [Cesa-Bianchi and Lugosi, 2006], and importance sampling. However, the dynamic setting we consider is fundamentally different: each round introduces a new piecewisestructured update to the weight function, and as these updates accumulate over t rounds in a d-dimensional domain, the number of distinct regions induced by the partition grows as O(td ) . This combinatorial explosion makes naively maintaining the weight function and sampling from it computationally infeasible—and yet, to our knowledge, the question of how to do this efficiently has received almost no attention. Prior work on online learning with piecewise-structured rewards [Balcan et al., 2018] has focused primarily on regret minimization and sample complexity, either assuming access to a sampling oracle or ignoring computational  \nefficiency altogether. ","cbCaibAdCG72yIHJ","https://ap.wps.com/l/cbCaibAdCG72yIHJ","pdf",832833,2,1,74,"English","en",105,"# Abstract\n# Introduction","[{\"question\":\"What problem does the document address?\",\"answer\":\"It addresses efficient online proportional sampling in a high-dimensional domain where the sampling distribution is defined by a dynamically evolving weight function induced by growing piecewise-structured partitions.\"},{\"question\":\"How does the work model the structure and changes over time?\",\"answer\":\"The domain is partitioned each round into regions with piecewise-polynomial reward functions whose boundaries are defined by hyperplanes from fixed directions (e.g., axis-parallel hyperplanes). Cumulative weights from these updates refine the partition over time.\"},{\"question\":\"What theoretical results and learning applications are provided?\",\"answer\":\"The document proves bounds on the depth of an efficient data structure under σ-smoothed adaptive adversaries (tight O(√σT)) and under random-order adversaries (O(log T)). It applies the framework to online learning with piecewise-structured rewards, obtaining efficient no-regret algorithms for both full-information and bandit feedback with sublinear regret.\"}]",1784202177,186,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"efficient-online-proportional-sampling-with-applications-to-smoothed-online-learning","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/efficient-online-proportional-sampling-with-applications-to-smoothed-online-learning/85268/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the document address?","Question",{"text":75,"@type":76},"It addresses efficient online proportional sampling in a high-dimensional domain where the sampling distribution is defined by a dynamically evolving weight function induced by growing piecewise-structured partitions.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the work model the structure and changes over time?",{"text":80,"@type":76},"The domain is partitioned each round into regions with piecewise-polynomial reward functions whose boundaries are defined by hyperplanes from fixed directions (e.g., axis-parallel hyperplanes). Cumulative weights from these updates refine the partition over time.",{"name":82,"@type":73,"acceptedAnswer":83},"What theoretical results and learning applications are provided?",{"text":84,"@type":76},"The document proves bounds on the depth of an efficient data structure under σ-smoothed adaptive adversaries (tight O(√σT)) and under random-order adversaries (O(log T)). It applies the framework to online learning with piecewise-structured rewards, obtaining efficient no-regret algorithms for both full-information and bandit feedback with sublinear regret.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]