[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83957-en":3,"doc-seo-83957-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83957,687197207639,"Asher","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","Optimism as a Vulnerability: Deceptive Stackelberg Control of UCB Bandit Followers","Upper Confidence Bound (UCB) learning ensures sublinear regret in stochastic bandits, but the optimism that drives statistical efficiency also creates a strategic vulnerability when interacting with an omniscient adaptive leader. The work formalizes the mismatch between classical strong Stackelberg equilibrium assumptions and a boundedly rational follower that learns from empirical reward histories. It proposes a finite-horizon deceptive leader mechanism with a honeypot signaling phase and a trap phase. Under explicit payoff and separation conditions, the leader’s cumulative utility exceeds the classical SSE ceiling, with manipulation cost bounded by O(√T ln T).","arXiv :2607 .05423v1 [ cs .GT] 28 Jun 2026  \nOptimism as a Vulnerability: Deceptive Stackelberg Control of UCB  \nBandit Followers  \n¸Suayp Talha Kocabay KOCABAYSUAYPTALHA08@GMAIL . COM  \nIndependent Researcher  \nKerem Yalçın 54 KEREMYALCIN @ GMAIL . COM  \nIndependent Researcher  \nTalha Rüzgar Akku¸s TALHARUZGARAKKUS@GMAIL . COM  \nIndependent Researcher  \nAbstract  \nUpper Confidence Bound (UCB) algorithms guarantee sublinear regret for agents learning unknown stochastic environments, yet the same principle that makes them statistically efficient—optimism in the face of uncertainty—induces a predictable strategic vulnerability against an omniscient adaptive leader. Classical strong Stackelberg equilibrium (SSE) assumes that the follower immediately best-responds to the leader’s committed mixed action; it therefore supplies no mechanism-design prescription for a leader facing a boundedly rational follower who constructsand acts on empirical reward histories. We formalize this conflict in a finite-horizon repeated Stackelberg game and give exact constructive proofs for a deceptive leader mechanism. In a honeypot phase, the leader pays a finite signaling cost to inflate the UCB index of a designated follower action. In a trap phase, the leader switches to a selfish action distribution while the follower remains locked into the designated action because the manipulated empirical history and exploration bonus dominate competing indices. Under explicit separation and payoff assumptions, the leader’s cumulative utility strictly exceeds the classical SSE ceiling, and the manipulation cost is bounded by a regret calculation of order O ( √T lnT) . The results identify a formal incompatibility between static equilibrium prescriptions and dynamically learned empirical incentives.  \nKeywords: Stackelberg games, bandit learning, UCB, strategic deception, reward manipulation  \n1. Introduction  \nStackelberg security games and related leader–follower models typically analyze commitment toa mixed action followed by a rational best response [5, 10, 21, 22] . In the strong Stackelberg equilibrium convention, the follower breaks ties in favor of the leader; the leader consequently solves a static optimization problem over induced best responses. This model is internally coherent when the follower observes payoffs and responds as a utility maximizer. It is not a model of a follower who learns payoffs from interaction.  \nBandit-learning followers instantiate a different behavioral primitive. UCB1 [1, 11, 20] is noregret in stationary stochastic bandits: arms with uncertain value receive an optimism bonus, and suboptimal arms are sampled only logarithmically often. In a strategic environment, however, reward samples are not exogenous evidence about a fixed arm. They are data produced by another player. This places the model closer to learning in games and nonstationary multi-agent learning  \n© ¸S.T. Kocabay, K. Yalçın & T.R. Akku¸s.  \nOPTIMISM AS A VULNERABILITY  \n[4, 6, 17, 19] than to one-shot commitment. An adaptive leader can therefore treat the follower’s statistical estimator as an object of control.  \nThe closest technical literature is reward or action poisoning of bandit learners, where an external attacker corrupts feedback or actions to force target pulls at small perturbation cost [2, 8, 12, 13, 15, 23, 26] . Our mechanism differs because the leader does not edit rewards exogenously; it creates the reward stream endogenously through legal Stackelberg play. Verification and corruptionaware bandits [9, 14, 18, 24] and Stackelberg learning with manipulative or non-myopic agents [3, 7, 16, 25] motivate the defensive discussion below.  \nThis paper makes the preceding claim precise. We compare two leader models in a repeated finite Stackelberg game. The baseline leader commits myopically to a classical SSE action and assumes immediate best response. The deceptive leader instead first rewards a target follower action j⋆ to raise its empirical mean","cbCaiq6njM60G6NT","https://ap.wps.com/l/cbCaiq6njM60G6NT","pdf",231471,5,1,13,"English","en",105,"# Abstract\n# Introduction\n# Model","[{\"question\":\"Why does UCB optimism become a strategic vulnerability in Stackelberg bandit settings?\",\"answer\":\"UCB’s exploration bonus and empirical mean can be exploited when the leader controls the reward stream endogenously through Stackelberg play, making the follower’s estimator a controllable object rather than passive evidence.\"},{\"question\":\"How does the deceptive leader mechanism work across phases?\",\"answer\":\"In a honeypot phase, the leader pays a finite signaling cost to inflate the UCB index of a designated follower action. In a trap phase, the leader changes to a selfish distribution while the follower keeps selecting the designated action because manipulated empirical history and the exploration bonus dominate alternatives.\"},{\"question\":\"What guarantees does the paper provide relative to the classical strong Stackelberg equilibrium ceiling?\",\"answer\":\"With separation, targetability, and exploitability assumptions, the deceptive leader achieves strictly higher cumulative utility than the classical SSE ceiling, while the manipulation cost is bounded by O(√T ln T).\"}]",1784191666,33,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"optimism-as-a-vulnerability-deceptive-stackelberg-control-of-ucb-bandit-followers","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/optimism-as-a-vulnerability-deceptive-stackelberg-control-of-ucb-bandit-followers/83957/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"Why does UCB optimism become a strategic vulnerability in Stackelberg bandit settings?","Question",{"text":76,"@type":77},"UCB’s exploration bonus and empirical mean can be exploited when the leader controls the reward stream endogenously through Stackelberg play, making the follower’s estimator a controllable object rather than passive evidence.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does the deceptive leader mechanism work across phases?",{"text":81,"@type":77},"In a honeypot phase, the leader pays a finite signaling cost to inflate the UCB index of a designated follower action. In a trap phase, the leader changes to a selfish distribution while the follower keeps selecting the designated action because manipulated empirical history and the exploration bonus dominate alternatives.",{"name":83,"@type":74,"acceptedAnswer":84},"What guarantees does the paper provide relative to the classical strong Stackelberg equilibrium ceiling?",{"text":85,"@type":77},"With separation, targetability, and exploitability assumptions, the deceptive leader achieves strictly higher cumulative utility than the classical SSE ceiling, while the manipulation cost is bounded by O(√T ln T).","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":20,"slug":138},19,"General","general"]