[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85597-en":3,"doc-seo-85597-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85597,1649267921044,"Ava Thompson","https://us-avatar.wpscdn.com/avatar/1800007509477c92dfb?_k=1782875107921204101",8,"Research & Report","Minimizing Worst-Case Weighted Latency for Multi-Robot Persistent Monitoring: Theory and RL-Based Solutions","Multi-robot persistent monitoring is investigated on weighted graphs where node weights represent monitoring priorities and edge weights represent travel distances. The objective is to minimize the worst-case weighted latency over all nodes across an infinite horizon, addressing shortcomings of standard worst-latency formulations that obscure transient weakness. A family of tail-performance objectives generalizes the classical metric, enabling theoretical guarantees, approximation via periodic solutions, and reductions to event-driven decision models.","Minimizing Worst-Case Weighted Latency for Multi-Robot Persistent Monitoring: Theory and RL-Based Solutions  \nJournal Title XX(X):1–39  \n©The Author(s) 2016  \nReprints and permission: [sagepub.co.uk/journalsPermissions.nav](sagepub.co.uk/journalsPermissions.nav)[ ](sagepub.co.uk/journalsPermissions.nav)DOI: 10.1177/ToBeAssigned [www.sagepub.com/](www.sagepub.com/)  \nSAGE  \narXiv :2605 .09633v2 [ cs .RO] 13 Jul 2026  \nWeizhen Wang1,2 , Ziheng Wang1 , Jianping He1,2 , Xinping Guan1,2 , and Xiaoming Duan1,2  \nAbstract  \nWe study multi-robot persistent monitoring on weighted graphs, where node weights encode monitoring priorities and edge weights encode travel distances. The goal is to design joint robot trajectories that minimize the worst-case weighted latency across all nodes over an infinite time horizon. The widely adopted worst-case latency objective evaluates team performance over the entire time horizon and therefore may fail to distinguish strategies with poor transient behavior but strong asymptotic performance. To address this limitation, we propose a family of tail-performance objectives that generalize the standard objective and study the resulting functional optimization problems. We establish several key theoretical properties, including the existence of optimal strategies, relationships among the proposed objectives and their corresponding optimization problems, approximation by periodic solutions to arbitrary accuracy, and reductions to event-driven decision models with discretized waiting time. Building on these results, we construct an equivalent event-driven Markov decision process (MDP), called the Tail Worst-case Latency-Optimizing Markov Decision Process (TWLO-MDP), which reformulates the tail-performance objective as a standard average-cost criterion. We then develop reinforcement-learning-based solution methods for the TWLO-MDP and introduce the multi-robot monitoring benchmark (M2Bench), a unified platform that supports the evaluation and comparison of heuristic and learning-based monitoring algorithms. Experiments on synthetic and realistic monitoring scenarios show that our methods effectively reduce the worst-case weighted latency and outperform representative baselines.  \nKeywords  \nMulti-robot monitoring, Markov decision process, reinforcement learning, benchmarking platform  \n1 Introduction  \nPersistent monitoring is a key application area for many real-world robotic systems, with representative scenarios including urban crime surveillance (PalmaBorda et al. 2026), post-disaster area monitoring (Kumar et al. 2022), anomaly detection (Witwicki et al. 2017), ocean monitoring (Smith et al. 2011), and forest fire detection (Momeni et al. 2022) . In these applications, a team of robots is required to repeatedly visit a set of important locations so that abnormal events can be detected in a timely manner. This naturally leads to a graph-based multi-robot monitoring problem, where nodes represent locations of interest, edges represent feasible travel paths, and the robots must coordinate their long-term movements on the graph over an infinite time horizon. A central performance measure for this problem is latency, which quantifies how long each location remains unvisited. Since large latency at any important location may delay event detection, many studies adopt a worst-case perspective and seek coordinated multirobot trajectories that minimize the maximum latency across the graph over time. This problem has attracted significant attention over the past two decades (Machado et al. 2003 ; Chevaleyre 2004 ; Portugal and Rocha 2011 ; Huang et al. 2019 ; Basilico 2022) . Earlier studies formalized the worst-latency objective and established the computational hardness of the resulting optimization problems, including  \nNP-hardness of exact optimization and APX-hardness of approximation in weighted variants (Chevaleyre 2004 ; Alamdari et al. 2014) . Subsequent work has developed a range of solution methods, incl","cbCaimyGSzyjM2jV","https://ap.wps.com/l/cbCaimyGSzyjM2jV","pdf",26405588,3,1,41,"English","en",105,"# Introduction\n# Problem formulation and tail-performance objectives\n# Tail Worst-case Latency-Optimizing MDP (TWLO-MDP)\n# Reinforcement-learning-based solutions\n# Benchmark M2Bench and experiments","[{\"question\":\"What is the monitoring objective in the document?\",\"answer\":\"The document targets joint multi-robot trajectories that minimize the worst-case weighted latency across all graph nodes over an infinite time horizon.\"},{\"question\":\"Why do the authors introduce tail-performance objectives instead of using the standard worst-case latency objective?\",\"answer\":\"They note that the standard objective may not distinguish strategies with poor transient behavior even if long-run (asymptotic) performance is strong.\"},{\"question\":\"How do reinforcement learning methods connect to TWLO-MDP in the solutions?\",\"answer\":\"The tail-performance objective is reformulated as an equivalent event-driven Markov decision process (TWLO-MDP), and reinforcement-learning-based solution methods are developed to solve this MDP formulation.\"}]",1784204832,103,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"minimizing-worst-case-weighted-latency-for-multi-robot-persistent-monitoring-theory-and-rl-based-solutions","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/minimizing-worst-case-weighted-latency-for-multi-robot-persistent-monitoring-theory-and-rl-based-solutions/85597/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is the monitoring objective in the document?","Question",{"text":75,"@type":76},"The document targets joint multi-robot trajectories that minimize the worst-case weighted latency across all graph nodes over an infinite time horizon.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Why do the authors introduce tail-performance objectives instead of using the standard worst-case latency objective?",{"text":80,"@type":76},"They note that the standard objective may not distinguish strategies with poor transient behavior even if long-run (asymptotic) performance is strong.",{"name":82,"@type":73,"acceptedAnswer":83},"How do reinforcement learning methods connect to TWLO-MDP in the solutions?",{"text":84,"@type":76},"The tail-performance objective is reformulated as an equivalent event-driven Markov decision process (TWLO-MDP), and reinforcement-learning-based solution methods are developed to solve this MDP formulation.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]