[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-128113-en":3,"doc-seo-128113-105":31,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":30},128113,3985741905716,"Rowan","https://ap-avatar.wpscdn.com/davatar_994ba38a5ba835b3df7d355c54d3ed8d",8,"Research & Report","Online Learning under Budget and ROI Constraints via Weak Adaptivity","Online learning is studied under budget and return-on-investment (ROI) long-term constraints, where a decision maker must select costly actions over T rounds to maximize expected reward. Standard constrained online algorithms assume advance knowledge of strict feasibility parameters (Slater parameters) and round-wise existence of strictly feasible offline optima. The work replaces these assumptions by equipping primal-dual templates with weakly adaptive regret minimizers, yielding a dual-balancing mechanism that keeps dual variables small without knowing Slater parameters. First best-of-both-worlds no-regret guarantees are proven for both stochastic and adversarial inputs, and instantiated for optimal bidding in mechanisms such as first-price auctions.","Online Learning under Budget and ROI Constraints via Weak Adaptivity  \nMatteo Castiglioni 1 Andrea Celli 2 Christian Kroer 3  \nAbstract  \nWe study online learning problems in which a decision maker has to make a sequence of costly decisions, with the goal of maximizing their expected reward while adhering to budget and return-on-investment (ROI) constraints. Existing primal-dual algorithms designed for constrained online learning problems under adversarial inputs rely on two fundamental assumptions. First, the decision maker must know beforehand the value of parameters related to the degree of strict feasibility of the problem (i.e. Slater parameters) . Second, a strictly feasible solution to the offline optimization problem must exist at each round. Both requirements are unrealistic for practical applications such as bidding in online ad auctions. In this paper, we show how such assumptions can be circumvented by endowing standard primal-dual templates with weakly adaptive regret minimizers. This results in a “dual-balancing” framework which ensures that dual variables stay sufficiently small, even in the absence of knowledge about Slater’s parameter. We prove the first best-ofboth-worlds no-regret guarantees which hold in absence of the two aforementioned assumptions, under stochastic and adversarial inputs. Finally, we show how to instantiate the framework to optimally bid in various mechanisms of practical relevance, such as first-price auctions.  \n1. Introduction  \nA decision maker takes decisions over T rounds. At each round t, the decision xt ∈ X is chosen before observing a reward function ft together with a set of time-varying  \n1DEIB, Politecnico di Milano, Milan, Italy 2Department of Computing Sciences, Bocconi University, Milan, Italy 3IEOR Department, Columbia University, New York, NY. Correspondence to: Matteo Castiglioni \u003C[matteo.castiglioni@polimi.it](matteo.castiglioni@polimi.it)>, Andrea Celli \u003C[andrea.celli2@unibocconi.it](andrea.celli2@unibocconi.it)>, Christian Kroer \u003Cchris[tian.kroer@columbia.edu](tian.kroer@columbia.edu)> .  \nProceedings of the 41 st International Conference on Machine Learning, Vienna, Austria. PMLR 235, 2024 . Copyright 2024 by the author(s) .  \nconstraint functions. The decision maker is allowed to make decisions that are not feasible, provided that the overall sequence of decisions obeys the long-term constraints over the entire time horizon, up to a small cumulative violation across the T rounds. The goal of the decision maker is to maximize their cumulative reward, while satisfying the longterm constraints. This model was first proposed by Mannor et al. (2009) and later developed along various directions (Mahdavi et al., 2012 ; Jenatton et al., 2016 ; Liakopouloset al., 2019 ; Yu et al., 2017 ; Castiglioni et al., 2022b) .  \nMotivated by applications in online ad auctions, we consider the case in which the decision maker has a budget anda return-on-investment (ROI) constraint (Auerbach et al., 2008 ; Golrezaei et al., 2023 ; 2021) . The decision maker is subject to bandit feedback: at each time t, the decision maker takes a decision xt and then observes the realized reward ft (xt) and a cost ct (xt) . Inputs (ft , ct) can either be generated i.i.d. or selected by an oblivious adversary.  \nA key challenge of our model is that ROI constraints are not packing, thereby preventing the direct application of known algorithms for adversarial bandits with knapsacks (ABwK)(Immorlica et al., 2022 ; Castiglioni et al., 2022a), or for online allocation problems with resource-consumption constraints (Balseiro et al., 2022) . Moreover, previous work addressing the adversarial case with non-packing constraints makes the assumption that the “worst-case feasibility” with respect to all constraint functions observed up to T is strictly positive (Sun et al., 2017 ; Castiglioni et al., 2022b ; Immorlica et al., 2022 ; Balseiro et al., 2022) . In other words, there has to exist a “safe” policy guarantee","cbCaijfFKevTThE8","https://ap.wps.com/l/cbCaijfFKevTThE8","pdf",488207,2,1,25,"English","en",105,"# Introduction\n## Contributions","[{\"question\":\"What problem does the paper study in online learning under budget and ROI constraints?\",\"answer\":\"It studies online decision making over T rounds that must maximize cumulative reward while satisfying long-term budget and ROI constraints.\"},{\"question\":\"Why do existing primal-dual constrained online learning algorithms face limitations in practice?\",\"answer\":\"They rely on knowing strict feasibility parameters in advance and on the existence of a strictly feasible offline solution at every round, which can fail in practical settings like ad auctions.\"},{\"question\":\"How does the proposed “dual-balancing” framework address these assumptions?\",\"answer\":\"It endows standard primal-dual templates with weakly adaptive regret minimizers, keeping dual variables sufficiently small even without knowledge of Slater’s parameter, and only requiring a safe policy frequently enough.\"}]","Online Learning under Budget and ROI Constraints via Weak Adaptivity | PDF",1785944903,63,{"code":4,"msg":32,"data":33},"ok",{"site_id":25,"language":24,"slug":34,"title":13,"keywords":35,"description":14,"schema_data":36,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":29},"online-learning-under-budget-and-roi-constraints-via-weak-adaptivity","",{"@graph":37,"@context":86},[38,54,69],{"@type":39,"itemListElement":40},"BreadcrumbList",[41,45,48,51],{"item":42,"name":43,"@type":44,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":46,"name":47,"@type":44,"position":20},"https://docshare.wps.com/document/","Document",{"item":49,"name":12,"@type":44,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":44,"position":53},"https://docshare.wps.com/document/online-learning-under-budget-and-roi-constraints-via-weak-adaptivity/128113/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":42,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-27","2026-08-05",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What problem does the paper study in online learning under budget and ROI constraints?","Question",{"text":76,"@type":77},"It studies online decision making over T rounds that must maximize cumulative reward while satisfying long-term budget and ROI constraints.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"Why do existing primal-dual constrained online learning algorithms face limitations in practice?",{"text":81,"@type":77},"They rely on knowing strict feasibility parameters in advance and on the existence of a strictly feasible offline solution at every round, which can fail in practical settings like ad auctions.",{"name":83,"@type":74,"acceptedAnswer":84},"How does the proposed “dual-balancing” framework address these assumptions?",{"text":85,"@type":77},"It endows standard primal-dual templates with weakly adaptive regret minimizers, keeping dual variables sufficiently small even without knowledge of Slater’s parameter, and only requiring a safe policy frequently enough.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":21,"doc_module":4,"doc_module_name":47,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":47,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":47,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":47,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":47,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":47,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":47,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":47,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":47,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":47,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":47,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]