[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-81622-en":3,"doc-seo-81622-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},81622,8796095360427,"Lucas Martin","https://ap-avatar.wpscdn.com/davatar_994ba38a5ba835b3df7d355c54d3ed8d",8,"Research & Report","Logarithmic High-Probability Regret for Strongly Convex Online Convex Optimization with Two-Point Bandit Feedback","Online convex optimization with two-point bandit feedback is studied against a nonanticipating adaptive adversary, where each round reveals only two function evaluations rather than gradients. For strongly convex losses, prior work established comparator-wise logarithmic regret in expectation, yielding pseudo-regret via minimizing outside the probability space. The main result proves a fixed-comparator high-probability two-query regret bound of logarithmic order in T and 1/δ, using high-confidence martingale-to-strong-convexity analysis and deterministic covering to extend to a realized full-comparator guarantee.","arXiv :2603 .25029v4 [ cs .LG] 9 Jul 2026  \nLogarithmic High-Probability Regret for Strongly Convex Online Convex Optimization with Two-Point Bandit Feedback  \nHaishan Ye  \nSchool of Management  \nXi’an Jiaotong University  \nhsye   [cs@outlook.com](cs@outlook.com)  \nAbstract  \nWe study online convex optimization (OCO) with two-point bandit feedback against a nonanticipating adaptive adversary. In this setting, a learner competes with an adversarial sequence of convex losses while observing each loss only through two function evaluations. For strongly convex losses, Agarwal, Dekel, and Xiao (2010) proved a comparator-wise logarithmic regret bound in expectation. Consequently, by minimizing outside the probability space, their result yields a pseudo-regret guarantee of the form EAT −minx∈K ELT (x), where AT is the algorithm’s two-query cumulative loss and LT (x) is the comparator’s cumulative loss. They asked whether a logarithmic high-probability guarantee is achievable in the same two-point strongly convex setting. Our main theorem provides the corresponding fixed-comparator high-probability statement: for any comparator x ∈ K fixed independently of the algorithmic random directions, the standard two-point projected gradient method guarantees, with probability at least 1 − δ, a two-query regret bound of order  \nO 􀀒 dGµ2 (log T + log(1/δ)) + dGD log(1/δ) + Glog T 􀀒 1 + Dr􀀓􀀓 .  \nAt the comparator-wise level, our leading horizon-dependent term is linear in d, compared with the d2-type term in the original analysis of Agarwal, Dekel, and Xiao. The key ingredient is a high-confidence analysis that simultaneously absorbs the martingale error into strong convexity and preserves the linear-in-dimension estimator control of the two-point method. A deterministic covering argument then yields a realized full-comparator guarantee against minx∈K LT (x), preserving logarithmic dependence on T at the cost of the standard covering-number factor.  \n1 Introduction  \nOnline convex optimization (OCO) is a repeated decision problem between a learner and an adversary. At each round t = 1 , ... , T, the learner chooses an action xt from a fixed convex set K ⊂ Rd. A non-anticipating adaptive adversary then selects a convex loss function ℓ t : K → R using the past history and, in the two-point protocol studied here, before the learner’s fresh random direction at round t is sampled. While full-information OCO is well understood, the bandit feedback setting, where only function values of ℓt are observed, remains substantially more challenging because gradients are unavailable. Two-point feedback is the minimal zeroth-order interface that permits an unbiased gradient estimate of a smoothed loss, making it a natural test case for whether bandit feedback can match full-information fast rates. Strong convexity is the canonical assumption under which full-information OCO improves from √T-type regret to logarithmic regret. High-probability guarantees are also important in online decision problems: an expected logarithmic bound may still allow rare but large deviations, whereas a high-confidence bound controls almost every realization of the learner’s randomization. The main difficulty is that the two-point gradient estimator is noisy, and this noise is coupled with the learner’s trajectory under an adaptive adversary. Consequently, the usual full-information logarithmic-regret proof and a direct Azuma-type concentration argument do not by themselves yield the desired rate.  \nFor compactness, write  \nT T  \nAT := Xt=1 ℓt(xt+~~ ~~αut)~~ ~~+2~~ ~~ℓt(xt~~ ~~−~~ ~~αut), LT (x) := Xt=1 ℓt(x) .  \nFor a fixed comparator x ∈ K, the two-point protocol analyzed in this paper measures performance by R2ptT(x) = AT−LT(x) . We distinguish three comparator semantics. A comparator-wise expected bound has the form  \n∀x ∈ K, E [AT − LT(x)] ≤ B.  \nSince x⋆exp ∈ arg minxELT(x) is deterministic, one may minimize outside the probability space and obtain the pseudo-regret bound  \nEAT ","cbCaigOM9nuS70VB","https://ap.wps.com/l/cbCaigOM9nuS70VB","pdf",372358,3,1,23,"English","en",105,"# Introduction\n## Problem setting and regret notions\n## Motivation from bandit feedback and strong convexity\n## Related work and missing ingredient","[{\"question\":\"What feedback model does the paper study in online convex optimization?\",\"answer\":\"It studies a two-point bandit protocol where the learner observes each convex loss only through two function evaluations (two-point queries) rather than full gradients.\"},{\"question\":\"What is the main theorem’s guarantee?\",\"answer\":\"For any fixed comparator x in the convex set K, the standard two-point projected gradient method achieves a two-query regret bound with probability at least 1−δ, with logarithmic dependence on T and 1/δ.\"},{\"question\":\"How does the result differ between comparator-wise and realized full-comparator regret?\",\"answer\":\"Comparator-wise statements hold for a fixed x and may depend on x through the high-probability event, while realized full-comparator regret requires a single guarantee over all x, obtained via a deterministic covering argument.\"}]",1784174887,58,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"logarithmic-high-probability-regret-for-strongly-convex-online-convex-optimization-with-two-point-bandit-feedback","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/logarithmic-high-probability-regret-for-strongly-convex-online-convex-optimization-with-two-point-bandit-feedback/81622/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What feedback model does the paper study in online convex optimization?","Question",{"text":75,"@type":76},"It studies a two-point bandit protocol where the learner observes each convex loss only through two function evaluations (two-point queries) rather than full gradients.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What is the main theorem’s guarantee?",{"text":80,"@type":76},"For any fixed comparator x in the convex set K, the standard two-point projected gradient method achieves a two-query regret bound with probability at least 1−δ, with logarithmic dependence on T and 1/δ.",{"name":82,"@type":73,"acceptedAnswer":83},"How does the result differ between comparator-wise and realized full-comparator regret?",{"text":84,"@type":76},"Comparator-wise statements hold for a fixed x and may depend on x through the high-probability event, while realized full-comparator regret requires a single guarantee over all x, obtained via a deterministic covering argument.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]