[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83219-en":3,"doc-seo-83219-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83219,962075114765,"Quinn","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","Nonlinear Bandit","This paper studies generalized linear bandits (GLB) under heavy-tailed noise, a setting widely seen in personalized recommendation, financial markets, and medical treatment. An online mirror descent (OMD) framework is used to propose EHM, extending an adaptive Huber loss with one-pass updates to achieve an almost optimal regret rate while keeping per-round computation constant in t and efficient in the horizon. A revised version, PGLB-EHM, preserves the regret order. The paper further develops NB-EHM for a nonlinear bandit special case using bisection and additional restrictions, and then applies affine lifting to extend to general nonlinear bandit problems with sublinear regret.","arXiv :2607 .07304v 1 [ cs .LG] 8 Jul 2026  \nNonlinear Bandit  \nTianshuo Zheng 1 , Ting Wu 1 , Zhi-Hua Zhou2 and Keqin Liu∗3  \n1 School of Mathematics, Nanjing University, Nanjing, 210093, China  \n2 School of Artificial Intelligence, Nanjing University, National Key Laboratory for Novel Software Technology, Nanjing, 210023, China  \n3 School of Mathematics and Physics, Xi’an Jiaotong-Liverpool University, Suzhou, 215123, China  \nAbstract  \nIn this paper we first study the problem of generalized linear bandit (GLB) under heavy-tailed noise. The characteristics of heavy-tailed distributions are widely observed in real-world applications such as personalized recommendation, financial markets, and medical treatments. Based on the online mirror descent (OMD) method, we propose an algorithm EHM that extends the adaptive Huber loss method [37] with one-pass update (O(1) computational complexity with respect to current round t and the time horizon T), which simultaneously achieves an almost optimal regret offust(iTot3G)2L]w,hoBeruperrTaoblgleisormthithue tmndiemerleimthhoine criatazoessentw.hhIeennnaecdeoddnittteiooxtnku,nabolywchutaarilcaizomctingmeriaostnislcpybecuseciaeoldmprpesoparpaiermectyetewoerisfesomNcoexnetst,liankwent, and we slightly revised former algorithm to obtain the PGLB-EHM algorithm. After theoretical analysis, we prove that the regret upper bound order stays the same. Furthermore, we look deeper into a special case of nonlinear bandit (NB) and present the NB-EHM algorithm with bisection method and special restriction. Eventually we utilize the affine lifting approach and show that the general NB problem can be applied with NB-EHM to achieve a sublinear regret bound.  \n1 Introduction  \nOnline sequential decision-making under uncertainty has long been a central challenge in reinforcement learning, operations research and stochastic optimization. How to balance between exploration and exploitation to maximize the long-term gained reward becomes the main question. The multi-armed bandit (MAB) has served as a canonical framework for studying sequential learning problems. By selecting an action in each round and receiving a random reward, the player aims to minimize the growth rate of regret (cost of learning), thereby achieving a higher reward gaining rate over the long run[22] . However, traditional MAB models assume a finite action set and ignore the underlying relationships between actions and the environmental context. Therefore, it limits bandit theory’s applicability in real-world scenarios such as online advertising and recommendation algorithms. The contextual bandit (CB) paradigm addresses this limitation by incorporating contextual features into the decision-making process. Within this paradigm, the linear bandit (LB) assumes that the expected reward depends linearly on multi-dimensional features, typically represented as the inner product of an action and a context vector. The modified model underpins classical algorithms such as LinUCB [24] and LinTS [1], which perform well in low-dimensional settings with light-tailed reward. However, many practical tasks, such as predicting user click-through rates, page views, and repeat purchase counts, exhibit highly non-linear relationships and complex, non-Gaussian reward distributions. Under such circumstances, simple linear models often fail to provide accurate predictions.  \nTo enhance the expressive power, generalized linear model (GLM) introduces a link function that maps linear predictions to a sample space better aligned with the data distribution, naturally accommodating distributions such as logistic, Poisson, and Pareto [29] . [14] were the first to integrate GLM into the bandit framework, proposing the GLM-UCB algorithm, while [25] provides a theoretical foundation for the generalized linear bandit (GLB) by establishing optimal regret bounds under high-dimensional features. Nevertheless, most existing GLB algorithms rely on maximum likelihood estimation (ML","cbCaipOWlCYcKugY","https://ap.wps.com/l/cbCaipOWlCYcKugY","pdf",749038,4,1,31,"English","en",105,"# Introduction\n## Bandit background: exploration-exploitation and regret\n## From MAB to contextual bandits\n## Linear bandits and limitations in non-linear settings\n## GLM/GLB and robustness challenges under noisy environments\n## Heavy-tailed noise and adaptive Huber loss\n## Nonlinear bandit: kernel, neural, and computational considerations","[{\"question\":\"What problem does the paper focus on?\",\"answer\":\"The paper focuses on generalized linear bandit learning under heavy-tailed noise, and extends the analysis to nonlinear bandit settings.\"},{\"question\":\"How does the proposed method achieve robustness to heavy-tailed noise?\",\"answer\":\"It is based on an adaptive Huber loss within an online mirror descent framework, combining L2 smoothness with L1 robustness to reduce bias from maximum-likelihood estimates.\"},{\"question\":\"What are the main algorithmic contributions and extensions?\",\"answer\":\"It introduces EHM with one-pass updates and derives a revised PGLB-EHM with the same regret order, then proposes NB-EHM for a nonlinear special case and extends to general nonlinear bandits via affine lifting for sublinear regret.\"}]",1784186020,78,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"nonlinear-bandit","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/nonlinear-bandit/83219/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the paper focus on?","Question",{"text":75,"@type":76},"The paper focuses on generalized linear bandit learning under heavy-tailed noise, and extends the analysis to nonlinear bandit settings.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the proposed method achieve robustness to heavy-tailed noise?",{"text":80,"@type":76},"It is based on an adaptive Huber loss within an online mirror descent framework, combining L2 smoothness with L1 robustness to reduce bias from maximum-likelihood estimates.",{"name":82,"@type":73,"acceptedAnswer":83},"What are the main algorithmic contributions and extensions?",{"text":84,"@type":76},"It introduces EHM with one-pass updates and derives a revised PGLB-EHM with the same regret order, then proposes NB-EHM for a nonlinear special case and extends to general nonlinear bandits via affine lifting for sublinear regret.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]