[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82044-en":3,"doc-seo-82044-105":29,"detail-sidebar-cat-0-en-105":83},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},82044,7971461740909,"Levi","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","DIRECTOR Accelerating Distributed MoE Serving via Online Proactive Expert Placement","Expert parallelism is the dominant paradigm for serving Mixture-of-Experts (MoE) models, where GPU communication and computation latencies are tightly coupled to how experts are placed. Prior expert-placement methods exploit historical activation patterns but degrade under diverse and rapidly changing request mixes. DIRECTOR presents an online proactive placement approach that predicts incoming expert activations, then migrates experts with near-zero downtime while addressing uncertainty, migration cost, and NP-hard optimization. A relaxation-based optimizer enforces capacity constraints with a (1+ε) approximation, and experiments show 11–55% end-to-end latency reduction across popular MoE models.","DIRECTOR: Accelerating Distributed MoE Serving via Online Proactive Expert Placement  \nQianli Liu 1 , Kaibin Guo2 , Zicong Hong 1 , Peng Li3 , Fahao Chen4 , Haodong Wang 1 , Jian Lin 1 , and Song Guo 1 1Department of Computer Science and Engineering, The Hong Kong University of Science and Technology, Hong Kong  \n2 School of Software Engineering, Sun Yat-Sen University, China  \n3 School of Cyber Science and Engineering, Xi’an Jiaotong University, China  \n4 School of Artificial Intelligence, Shandong University, China  \n[qianli.liu@connect.ust.hk](qianli.liu@connect.ust.hk), [guokb@mail2.sysu.edu.cn](guokb@mail2.sysu.edu.cn), [ziconghong@gmail.com](ziconghong@gmail.com), [pengli@xjtu.edu.cn](pengli@xjtu.edu.cn),  \n[chenfh@ieee.org](chenfh@ieee.org), [hwanghb@connect.ust.hk](hwanghb@connect.ust.hk), [jlindc@connect.ust.hk](jlindc@connect.ust.hk), [songguo@cse.ust.hk](songguo@cse.ust.hk)  \narXiv :2607 .08782v1 [ cs .LG] 13 Jun 2026  \nAbstract—Expert parallelism has become the prevailing paradigm to serve Mixture-of-Experts (MoE) models. Its efficiency depends on the communication and computation latencies of the GPUs, which are linked to the placement of experts in the GPUs. Existing works for optimizing expert placement focus on leveraging past requests’ expert activation patterns. However, they demonstrate deficiencies facing diverse and rapidly changing request patterns, calling for an online, proactive approach. Implementing such an approach requires addressing several challenges: the uncertainty associated withincoming requests’ expert activation, the cost of expert migration, and the NP-hard complexity in optimization. Therefore, we present DIRECTOR, a new distributed MoE serving system that minimizes end-to-end latency via prediction-driven, online expert placement. DIRECTOR uses either a lightweight cascaded predictor or a low-bit quantized replica for expert activation patterns of incoming requests. An online migration module then enacts the changes with near-zero downtime by executing migrations in compute-bound phases, keeping disruption bounded. At its core, a relaxation-based expert placement optimizer operates under capacity constraints, runs in polynomial time, and achieves a (1 + ϵ) approximation ratio. Finally, we implement a prototype and demonstrate, through extensive experiments, a reduction in end-to-end latency of 11 ∼ 55% for popular MoE models (e.g., Mistral, DeepSeek and Qwen) compared to existing work.  \nI. INTRODUCTION  \nMixture-of-Experts (MoE) models [1]–[7] have become a popular architecture for scaling large language models (LLMs) to trillions of parameters while keeping the computational cost proportional. Unlike dense models, which use all parameters, MoE models dynamically route input tokens to a subset of feed-forward networks (FFNs), known as experts. This selective activation allows the required FLOPs to scale sublinearly as the model size increases [8]–[10] .  \nThe increasing parameter count of MoE models requires expert parallelism for distributed MoE serving due to the  \nThis research was supported by fundings from the Hong Kong RGC General Research Fund (152169/22E, 152228/23E, 162161/24E, 162116/25E), Research Impact Fund (No. R5060-19, No. R5011-23), Collaborative Research Fund (No. C1042-23GF), NSFC/RGC Collaborative Research Scheme (Grant No. 62461160332 & CRS HKUST602/24), Areas of Excellence Scheme (AoE/E-601/22-R), National Natural Science Foundation of China (No. 62471383), and the InnoHK (HKGAI) . Corresponding authors: Zicong Hong, Song Guo.  \nMoE Model  \nExpert Placement Strategy  \nOffline  \nHistory Requests  \nOnline Online  \nReactive Proactive (Ours)  \nThe Last Batch  \nRequests in Queue  \nStart Now  \nFig. 1: A comparison of offline expert placement, online reactive expert placement, and online proactive expert placement.  \nsignificant memory footprint [11] . In this paradigm, the experts are distributed across a cluster of GPUs, while non-expert parameters, such as self-attenti","cbCaifrvZnhhizhI","https://ap.wps.com/l/cbCaifrvZnhhizhI","pdf",868521,1,10,"English","en",105,"# Introduction\n## Overview of Mixture-of-Experts Serving\n## Bottlenecks in Distributed MoE\n## Related Work: Offline vs Online Reactive Placement\n## DIRECTOR: Online Proactive Placement","[{\"question\":\"What guarantees does the placement optimizer provide?\",\"answer\":\"DIRECTOR’s relaxation-based optimizer runs in polynomial time under capacity constraints and achieves a (1+ε) approximation ratio.\"}]",1784177788,25,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":78,"head_meta":80,"extra_data":82,"updated_unix":27},"director-accelerating-distributed-moe-serving-via-online-proactive-expert-placement","",{"@graph":35,"@context":77},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/director-accelerating-distributed-moe-serving-via-online-proactive-expert-placement/82044/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71],{"name":72,"@type":73,"acceptedAnswer":74},"What guarantees does the placement optimizer provide?","Question",{"text":75,"@type":76},"DIRECTOR’s relaxation-based optimizer runs in polynomial time under capacity constraints and achieves a (1+ε) approximation ratio.","Answer","https://schema.org",{"og:url":51,"og:type":79,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":81,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":84},[85,89,93,97,102,107,112,115,120,123,126],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":86,"show_sort_weight":87,"slug":88},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":90,"show_sort_weight":91,"slug":92},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Exam",70,"exam",{"id":98,"doc_module":4,"doc_module_name":45,"category_name":99,"show_sort_weight":100,"slug":101},5,"Comic",60,"comic",{"id":103,"doc_module":4,"doc_module_name":45,"category_name":104,"show_sort_weight":105,"slug":106},6,"Technology",50,"technology",{"id":108,"doc_module":4,"doc_module_name":45,"category_name":109,"show_sort_weight":110,"slug":111},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":113,"slug":114},30,"research-report",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},9,"Religion & Spirituality",20,"religion-spirituality",{"id":118,"doc_module":4,"doc_module_name":45,"category_name":121,"show_sort_weight":118,"slug":122},"World Cup","world-cup",{"id":21,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":21,"slug":125},"Lifestyle","lifestyle",{"id":127,"doc_module":4,"doc_module_name":45,"category_name":128,"show_sort_weight":98,"slug":129},19,"General","general"]