[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85152-en":3,"doc-seo-85152-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85152,1374391974468,"Eden","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","Distributionally Robust Multi-agent Reinforcement Learning Framework for Intelligent Intersection Control","Multi-agent reinforcement learning for traffic signal control often optimizes expected returns under nominal conditions, making learned policies vulnerable to spatial–temporal demand shifts and worst-case congestion. The paper presents an algorithm-agnostic Distributionally Robust MARL framework that couples an adaptive Contextual Bandit Worst-Case Estimator with traffic controllers during training to generate adversarial demand mixtures. Evaluations across multiple MARL families on grid and Monaco City networks prevent unbounded queue growth, significantly improving worst-case robustness and average-case efficiency, including strong zero-shot generalization.","A Distributionally Robust Multi-agent Reinforcement Learning Framework for Intelligent Intersection Control  \nShuwei Pei, Joran Borger, Arda Kosay, Bayu Jayawardhana, Muhammed O. Sayin, Saeed Ahmed  \narXiv :2607 .09899v1 [ ee ss . SY] 10 Jul 2026  \nAbstract—Multi-agent reinforcement learning (MARL) has emerged as a promising approach for traffic signal control. However, standard MARL policies typically optimize for expected returns under nominal conditions, leaving them highly vulnerable to spatial-temporal demand shifts and catastrophic congestion under adverse scenarios. To address this critical limitation, this paper proposes an algorithm-agnostic Distributionally Robust (DR) MARL framework integrating an adaptive ContextualBandit Worst-Case Estimator (CB-WCE). Operating on a slower timescale, the CB-WCE co-evolves with the traffic controllers by dynamically generating adversarial demand mixtures during training. This steers the learning process to fortify policies against bottleneck scenarios without requiring modifications to the underlying MARL architectures. The framework is evaluated across value-based, actor-critic, and policy-gradient methods on both a synthetic 5 × 5 grid and a heterogeneous Monaco City network. Empirical results demonstrate that the DR framework prevents unbounded queue growth and profoundly enhances both worst-case robustness and average-case efficiency. Notably, for the Proximal Policy Optimization (PPO) architecture in the Monaco environment, on average, robust retraining reduced the worstcase queue length by 74.39% and improved the average-case network-wide queue length by 75.45% . Furthermore, the retrained policies exhibit strong zero-shot generalization to unseen traffic distributions, highlighting the framework’s scalability and potential for resilient real-world urban deployment.  \nIndex Terms—Reinforcement Learning; Distributionally Robust Optimization; Traffic Signal Control; Intelligent Transportation Systems.  \nI. INTRODUCTION  \nSignalized intersections are critical bottlenecks in urban road networks, contributing significantly to travel delays, fuel consumption, and pollutant emissions. Suboptimal signal timing exacerbates congestion, causing economic and environmental impacts [1], while also worsening air quality and associated public health risks [2] . As urbanization and population growth increasingly strain saturated transport infrastructure [3], developing traffic signal control strategies resilient to highly variable and uncertain demand remains a paramount challenge for sustainable urban mobility.  \nConventional traffic signal control predominantly relies on fixed-time plans, actuated logic, or rule-based adaptive schemes. Widely deployed systems, such as SCOOT [4] and SCATS [5], adapt cycle lengths and splits using historical data  \nThis work was supported by the Holland High Tech (TKI HTSM) strategic program PPS-I Flex HighTech under the project number 24PPS173-CABS. Shuwei Pei, Joran Borger, Bayu Jayawardhana, and Saeed Ahmed are with Engineering and Technology Institute Groningen, Faculty of Science and Engineering, University of Groningen, 9747 AG Groningen, the Netherlands. Arda Kosay and Muhammed O. Sayin are with the Department of Electrical, Electronics Engineering, Bilkent University, TR-06800 Ankara, Turkey. Email: [s.pei@rug.nl](s.pei@rug.nl), [j.borger.3@student.rug.nl](j.borger.3@student.rug.nl), [arda.kosay@bilkent.edu.tr](arda.kosay@bilkent.edu.tr), [b.jayawardhana@rug.nl](b.jayawardhana@rug.nl), [sayin@ee.bilkent.edu.tr](sayin@ee.bilkent.edu.tr), [s.ahmed@rug.nl](s.ahmed@rug.nl)  \nFig. 1. Schematic of the proposed DR-MARL training framework for intelligent intersection management. The diagram is adapted and extended from our previous work [9] to reflect the broader set of traffic networks and learning algorithms considered in this paper.  \nand localized real-time measurements. Although optimizationbased approaches like OPAC [6] and PRODYN [7] formulate signal c","cbCain4ksyKHJbkb","https://ap.wps.com/l/cbCain4ksyKHJbkb","pdf",8923601,3,1,14,"English","en",105,"# Introduction\n## Problem of nominal-trained MARL under demand shifts\n## Motivation for robustness and tail-performance guarantees\n# Proposed DR-MARL Framework\n## Adaptive Contextual Bandit Worst-Case Estimator\n## Algorithm-agnostic integration with MARL architectures\n# Experimental Evaluation\n## Grid environment results\n## Monaco City network results\n# Conclusion","[{\"question\":\"What problem does the paper address in existing multi-agent reinforcement learning for traffic signals?\",\"answer\":\"It addresses the vulnerability of MARL policies trained on nominal demand distributions, which can lead to severe congestion and poor tail performance under adverse spatial–temporal demand shifts.\"},{\"question\":\"How does the proposed DR-MARL framework improve robustness?\",\"answer\":\"It introduces an algorithm-agnostic distributionally robust objective with an adaptive Contextual Bandit Worst-Case Estimator that generates adversarial demand mixtures during training.\"},{\"question\":\"What evidence of performance gains does the paper report on the Monaco environment?\",\"answer\":\"For PPO, robust retraining reduced the worst-case queue length by 74.39% and improved the average-case network-wide queue length by 75.45%, while also demonstrating strong zero-shot generalization to unseen traffic distributions.\"}]",1784201416,35,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"distributionally-robust-multi-agent-reinforcement-learning-framework-for-intelligent-intersection-control","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/distributionally-robust-multi-agent-reinforcement-learning-framework-for-intelligent-intersection-control/85152/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-22","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the paper address in existing multi-agent reinforcement learning for traffic signals?","Question",{"text":75,"@type":76},"It addresses the vulnerability of MARL policies trained on nominal demand distributions, which can lead to severe congestion and poor tail performance under adverse spatial–temporal demand shifts.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the proposed DR-MARL framework improve robustness?",{"text":80,"@type":76},"It introduces an algorithm-agnostic distributionally robust objective with an adaptive Contextual Bandit Worst-Case Estimator that generates adversarial demand mixtures during training.",{"name":82,"@type":73,"acceptedAnswer":83},"What evidence of performance gains does the paper report on the Monaco environment?",{"text":84,"@type":76},"For PPO, robust retraining reduced the worst-case queue length by 74.39% and improved the average-case network-wide queue length by 75.45%, while also demonstrating strong zero-shot generalization to unseen traffic distributions.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]