[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82119-en":3,"doc-seo-82119-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},82119,1099514067415,"Rowan","https://ap-avatar.wpscdn.com/avatar/100002539d78ffe74a7?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779092875211072502",8,"Research & Report","Stochastic Linear Bandits with Partially Observed Actions","The stochastic linear bandit models sequential decision-making where each action is a feature vector and expected reward is linear in unknown parameters. This work studies a partially observed variant: for every action, the learner observes only a random subset of coordinates, reflecting realistic constraints such as missing item features in recommendation and data-collection limits in healthcare. Sublinear regret is generally information-theoretically impossible, but low intrinsic dimension of actions enables it. The TOFU-POV method estimates a latent action subspace from masked observations, imputes actions via an epoch-frozen representation, and applies OFUL in the learned low-dimensional coordinates. Theoretical guarantees deliver √T regret depending on intrinsic dimension and quantify effects of missingness, decision set size, and subspace conditioning, supported by rank-adaptive variants, lower bounds, and experiments on synthetic and real data.","Stochastic Linear Bandits with Partially Observed Actions ∗  \nGautam Dasarathy [gautamd@asu.edu](gautamd@asu.edu)  \nArizona State University  \nVineet Gattani [gattanivineet@gmail.com](gattanivineet@gmail.com)  \nGE Vernova  \nLalit Jain [lalitkumarj@gmail.com](lalitkumarj@gmail.com)  \nGoogle  \narXiv :2607 .0897 1v 1 [ cs .LG] 9 Jul 2026  \nAbstract  \nThe stochastic linear bandit, where actions are represented as vectors and rewards are linear, is a central paradigm for sequential decision making. We study a partially observed variant of this problem in which the learning agent only sees a random subset of coordinates for each action. Such partial observability arises naturally in settings like recommendation and healthcare, where full action descriptions can be expensive or even impossible to obtain. In general, this makes sublinear regret information-theoretically impossible. However, we show that this barrier can be overcome when the action vectors have low intrinsic dimension. We propose an algorithm, TOFU-POV, that estimates the latent action subspace using the masked actions, imputes current actions using an epoch-wise frozen representation, and runs OFUL in the resulting low-dimensional coordinates. Our theory shows that TOFU-POV enjoys a √T regret that scales with the intrinsic action subspace dimension as opposed to the ambient dimension and quantifies the interaction between these quantities and the missingness, decision set size, and subspace conditioning. We also devise a rank-adaptive algorithm that does not require the knowledge of the intrinsic dimension. We complement these guarantees with a lower bound based on a novel product construction that separates usual reward-learning uncertainty from a missingness-dependent cost intrinsic to partial observation. Synthetic and real data experiments support our theory and show that TOFU-POV can substantially improve upon natural baselines in this challenging problem.  \n1 Introduction  \nThe stochastic linear bandit (SLB) is an important framework for sequential decision-making under uncertainty, where the expected reward of a vector-valued action is assumed to be a linear function of its features [1–3] . At each round t, the learning agent is presented with a decision set Dt = {Xt,1, Xt,2 , . . .} and chooses an action Xt ∈ Dt, which results in a reward rt = ⟨Xt, θ⋆ ⟩ + η t , where θ⋆ ∈ Rd is unknown and η t is conditionally zero-mean random noise. The goal here is to minimize the cumulative regret relative to an oracle that (a) knows θ⋆ and, therefore, (b) chooses the action in Dt that maximizes the expected reward each round. A widely studied algorithm in this  \n∗This work was primarily performed while Vineet Gattani was a PhD student at Arizona State University, and was partially supported by the National Science Foundation award CCF-2048223 .  \nsebtotingundsisuOFp tUoLlog[ 1a]r,iwhose regretthmic factoriss knowSLBsnhto baveefboundedound far-abranovgienbgya(dliTti,mnsaitcnhriencgokmnomwnendloatwioern systems, advertising, and treatment allocation [4–6] .  \nIn many modern applications, however, observing the full feature vector of each action is prohibitively expensive, infeasible, or impossible. In recommendation systems [4], due to privacy, storage, or computational constraints, only a sparse subset of item features may be accessible. Similarly, in scientific or healthcare applications, constraints on sensing or data collection may naturally lead to missing observations. This motivates a more challenging variant of the SLB problem, where at each round, the agent only observes a subset of entries from each action vector—what we refer to as partial observability. Formally, for each action vector Xt ∈ Rd, each coordinate is revealed independently with probability p ∈ (0 , 1] . Without any further structure, reward-learning is information-theoretically impossible here: the agent cannot infer the full linear reward model, and pays a suboptimality price (i.e., regret) that is ","cbCaigOTiBwW4aMo","https://ap.wps.com/l/cbCaigOTiBwW4aMo","pdf",1786674,1,54,"English","en",105,"# Abstract\n# Introduction","[{\"question\":\"What is the partially observed stochastic linear bandit problem studied here?\",\"answer\":\"At each round, the agent selects an action from a decision set, but only a random subset of that action’s coordinates is revealed. Rewards still depend linearly on the full action vector through an unknown parameter vector.\"},{\"question\":\"Why is sublinear regret impossible in general, and what structural assumption changes this?\",\"answer\":\"Without additional structure, the learner cannot identify the full linear reward model from incomplete observations, leading to regret that grows linearly with time. When action vectors have low intrinsic dimension (lie near an unknown low-dimensional subspace), sublinear regret becomes achievable.\"},{\"question\":\"How does the proposed TOFU-POV algorithm overcome partial observability?\",\"answer\":\"TOFU-POV estimates the latent action subspace using masked coordinates, imputes current actions using an epoch-wise frozen subspace representation, and then runs OFUL in the resulting low-dimensional coordinates to control regret.\"}]",1784178315,136,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"stochastic-linear-bandits-with-partially-observed-actions","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/stochastic-linear-bandits-with-partially-observed-actions/82119/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is the partially observed stochastic linear bandit problem studied here?","Question",{"text":75,"@type":76},"At each round, the agent selects an action from a decision set, but only a random subset of that action’s coordinates is revealed. Rewards still depend linearly on the full action vector through an unknown parameter vector.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Why is sublinear regret impossible in general, and what structural assumption changes this?",{"text":80,"@type":76},"Without additional structure, the learner cannot identify the full linear reward model from incomplete observations, leading to regret that grows linearly with time. When action vectors have low intrinsic dimension (lie near an unknown low-dimensional subspace), sublinear regret becomes achievable.",{"name":82,"@type":73,"acceptedAnswer":83},"How does the proposed TOFU-POV algorithm overcome partial observability?",{"text":84,"@type":76},"TOFU-POV estimates the latent action subspace using masked coordinates, imputes current actions using an epoch-wise frozen subspace representation, and then runs OFUL in the resulting low-dimensional coordinates to control regret.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]