[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85224-en":3,"doc-seo-85224-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85224,13056703019662,"Evangeline","https://ap-avatar.wpscdn.com/avatar/be000253a8e92610077?_k=1778726343310543188",8,"Research & Report","Learning from Local Walks on Dynamic Graphs with Bandit Feedback","Stochastic multi-armed bandits are studied on dynamic graphs where arms are vertices and edges change over time. The learner must move locally, choosing either the current node or an immediate neighbor each round, so best-arm identification does not guarantee reachability for exploitation. A process-agnostic structural requirement based on sliding-window mixing keeps the graph’s intrinsic walk stable for exploration and navigation. Under this regime, local explore-then-commit algorithms achieve sublinear expected regret, including a reward-aware strategy with safety and performance-gain theorems.","Learning from Local Walks on Dynamic Graphs with Bandit Feedback  \nSourav Chakraborty∗ , Amit Kiran Rege∗ , Claire Monteleoni, Lijun Chen  \narXiv :2607 . 1057 1v 1 [ cs .LG] 12 Jul 2026  \nAbstract—We study stochastic multi-armed bandits on dynamic graphs, where arms correspond to the vertices of a network with time-varying edges. In this setting, the learner is restricted to local movement, selecting only its current node or an immediate neighbor at each round. This constraint decouples best-arm identification from exploitation: even after the optimal arm is identified, the learner may remain unable to reach it through the evolving topology. We identify a process-agnostic structural condition, based on sliding-window mixing, that ensures the graph’s intrinsic walk remains stable for both exploration and navigation. Under this regime, we analyze a family of local explore-then-commit algorithms and establish sublinear expected regret. Our framework includes a rewardaware strategy, for which we prove a worst-case safety theorem and a separate performance gain theorem.  \nI. INTRODUCTION  \nIn the stochastic multi-armed bandit (MAB) problem, a learner repeatedly chooses among several actions (interchangeably referred to as arms) observes only the reward of the selected action, and seeks to maximize its cumulative reward ([2], [9], [16]) . Classical bandit models typically assume that every arm is available for selection at every round. This assumption, however, is too strong for many networked settings where movement is restricted and inherently local. For instance, a mobile data collector in a sensor network is constrained by physical proximity, requiring it to be near a specific node to harvest its data [10]; a search routine on a peer-to-peer overlay is confined to traversing logical links between neighboring peers [8]; and a patrol agent is restricted to navigable pathways that may open or close asthe environment changes [4] . In these scenarios, the learner cannot jump to an arbitrary location; instead, its next action is strictly determined by its current position and the connectivity available at that moment.  \nThis motivates the dynamic graph bandit problem, where arms correspond to vertices in a network with time-varying edges. The learner observes only its local neighborhood, moves along locally available edges, and receives feedback solely from visited nodes. Consequently, the evolving topology governs both immediate reward availability and future reachability. This fundamentally differs from feedbackgraph bandits, where graphs encode side-observations rather than movement constraints ([11], [1]), and from static-graph bandits, where fixed, globally known topologies permit preplanned navigation ([18], [13]) .  \n∗Equal Contribution  \nAll authors are with the Department of Computer Science, University of Colorado, Boulder, CO 80309, USA { sourav .chakraborty,  \namit.rege, cmontel, [lijun.chen](lijun.chen}@colorado.edu)[}](lijun.chen}@colorado.edu)[@colorado.edu](lijun.chen}@colorado.edu)[ ](lijun.chen}@colorado.edu)Claire Monteleoni is also with INRIA, Paris, France.  \nA core challenge in this setting is that best-arm identification is decoupled from exploitation. In classical bandits, identifying the optimal arm immediately enables perpetual exploitation. Here, a learner might identify the best arm but remain temporarily unable to reach it. Learnability thus requires the environment to consistently provide navigable pathways; otherwise, the learner could be structurally isolated indefinitely. Importantly, standard full-horizon conditions, such as the connectivity of the union of all graphs, are insufficient, as a network might be connected in aggregate but disconnected at almost every individual step. A valid structural condition must therefore be shift-invariant, holding from any arbitrary round. This guarantees navigability during the post-exploration phase, regardless of the random stopping time at which exploration con","cbCaindk9HdbzxLP","https://ap.wps.com/l/cbCaindk9HdbzxLP","pdf",1093365,4,1,23,"English","en",105,"# Introduction\n## Dynamic graph bandits and local movement constraints\n## Decoupling best-arm identification from exploitation\n## Explore-then-commit framework and canonical walk\n## Common-stationary sliding-window mixing condition","[{\"question\":\"How does the dynamic graph bandit setting differ from classical bandits?\",\"answer\":\"Classical bandits assume every arm is available at every round, while here arms are vertices in a network with time-varying edges. The learner can only move locally by staying at the current node or moving to an immediate neighbor, and it receives feedback only from visited nodes.\"},{\"question\":\"Why is best-arm identification decoupled from exploitation in this problem?\",\"answer\":\"Even if the learner identifies the optimal arm, evolving topology may prevent reaching it during the exploitation phase. Learnability requires that navigable pathways consistently exist after exploration.\"},{\"question\":\"What structural condition ensures stable learning and navigation?\",\"answer\":\"The document introduces a process-agnostic common-stationary sliding-window mixing property. It requires the canonical walk to maintain a consistent long-run stationary distribution and enough well-connected rounds in every short window to avoid getting trapped in a local subgraph.\"}]",1784201848,58,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"learning-from-local-walks-on-dynamic-graphs-with-bandit-feedback","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/learning-from-local-walks-on-dynamic-graphs-with-bandit-feedback/85224/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"How does the dynamic graph bandit setting differ from classical bandits?","Question",{"text":75,"@type":76},"Classical bandits assume every arm is available at every round, while here arms are vertices in a network with time-varying edges. The learner can only move locally by staying at the current node or moving to an immediate neighbor, and it receives feedback only from visited nodes.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Why is best-arm identification decoupled from exploitation in this problem?",{"text":80,"@type":76},"Even if the learner identifies the optimal arm, evolving topology may prevent reaching it during the exploitation phase. Learnability requires that navigable pathways consistently exist after exploration.",{"name":82,"@type":73,"acceptedAnswer":83},"What structural condition ensures stable learning and navigation?",{"text":84,"@type":76},"The document introduces a process-agnostic common-stationary sliding-window mixing property. It requires the canonical walk to maintain a consistent long-run stationary distribution and enough well-connected rounds in every short window to avoid getting trapped in a local subgraph.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]