[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82909-en":3,"doc-seo-82909-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},82909,8796095461564,"Liam","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","Communication-Aware Placement and Pruning for Efficient Mixture-of-Experts Inference","As Mixture-of-Experts (MoE) models scale to hundreds of experts, expert placement and pruning increasingly control the communication volume, which directly impacts distributed inference performance across GPUs and nodes. This work introduces CAP (Communication-Aware Assignment and Pruning), integrating computation, communication, and accuracy. CAP uses co-activation-driven expert placement to group frequently co-activated experts, adjusts placements to balance communication versus load, and performs communication-aware pruning to reduce routing destinations with limited accuracy loss. Experiments on single-node and multi-node settings show 1.23×–1.86× throughput gains while preserving accuracy under similar speedups.","Communication-Aware Placement and Pruning for Efficient Mixture-of-Experts Inference  \nXiao Shi, Yingying Sun, Jiangsu Du, Zhiguang Chen, Yutong Lu  \nSun Yat-sen University  \nGuangzhou, China  \nEmail: {shix36, [sunyy57](sunyy57}@mail2.sysu.edu.cn)[}](sunyy57}@mail2.sysu.edu.cn)[@mail2.sysu.edu.cn](sunyy57}@mail2.sysu.edu.cn), {dujiangsu, chenzhg29, [luyutong](luyutong}@mail.sysu.edu.cn)[}](luyutong}@mail.sysu.edu.cn)[@mail.sysu.edu.cn](luyutong}@mail.sysu.edu.cn)  \narXiv :2607 .05 1 16v 1 [ cs .DC] 6 Jul 2026  \nAbstract—As MoE models scale to hundreds of experts, placement and pruning decisions increasingly dictate communication volume, affecting the performance of distributed inference across GPUs and nodes. We propose CAP (CommunicationAware Assignment and Pruning), a framework that considers computation, communication and accuracy together for efficient MoE inference through expert placement and pruning. It consists of three components: (1) Co-activation driven expert placement, which groups frequently co-activated experts to reduce inter-device and inter-node communication; (2) Communicationcomputation trade-off adjustment, which generates placements with different computational load and communication volume; and (3) Communication-aware expert pruning, which selectively removes routing destinations to reduce communication with limited accuracy degradation. By combining these components, CAP selects an efficient operating strategy for different hardware configurations. Across our single-node and multi-node experiments, it achieves 1.23×–1.86× throughput improvement over DeepSeek EPLB and sequential placement in vLLM, and preserves better model accuracy at the same target speedup under lossy acceleration.  \nIndex Terms—LLM, MoE inference, Expert Parallelism  \nI. INTRODUCTION  \nThe sparsely activated Mixture-of-Experts (MoE) architecture is increasingly used to scale the size of large language model (LLM) and boost performance [1]–[6] . A key challenge in MoE inference with expert parallelism is the imbalance in workload distribution across devices. Many efforts are devoted to adjusting expert placement and pruning for better performance. However, as the number of experts increases, existing approaches generally overlook a critical factor: the fact that expert placement and pruning significantly affect communication volume, leading to sub-optimal performance. MoE inference systems typically adopt a combination of expert and data parallelism. As shown in Figure 1, the data parallelism is applied to the dense attention part, where each device replicates the full weights of the dense layers, while expert parallelism is applied to the expert parts, where experts are partitioned by expert and distributed across all devices.  \nWhen input sequences arrive, they are processed in different devices separately, and then each token is routed to the devices hosting its target experts through a global all-to-all communication. The token is then returned to the originating device through another global all-to-all communication, so that subsequent attention computation of sequences can be performed.  \nInput Tokens  \nDense Part Expert Part  \nFig. 1: Expert parallelism combined with data parallelism  \nExpert placement and expert pruning have become common approaches for optimizing MoE inference systems. Expert placement methods [7]–[12] are lossless. Since the activation frequencies of different experts remain uneven, the distributed MoE inference generally suffers from severe load imbalance. By co-locating frequently and infrequently activated experts on the same device, it can balance the computation load and improve overall inference efficiency. Expert pruning methods [13]–[16] are typically lossy. They dynamically remove negligible experts for each token and trade a small amount of accuracy for higher inference speed.  \nSince prior MoE models featured only a small number of experts, existing approaches overlooked their impacts on communica","cbCailX9uMmumzBo","https://ap.wps.com/l/cbCailX9uMmumzBo","pdf",2184963,4,1,12,"English","en",105,"# Introduction\n## Background: Expert parallelism in MoE\n## Prior work: placement and pruning trade-offs\n## Proposed approach: CAP","[{\"question\":\"What problem does CAP address in MoE inference?\",\"answer\":\"CAP targets sub-optimal MoE inference caused by expert placement and pruning decisions that ignore their effect on communication volume, which becomes a dominant bottleneck at large expert counts.\"},{\"question\":\"How does CAP determine expert placement?\",\"answer\":\"CAP uses co-activation-driven placement to group frequently co-activated experts, then generates candidate placements by adjusting the communication–computation trade-off to explore different load and communication balances.\"},{\"question\":\"What is communication-aware expert pruning in CAP?\",\"answer\":\"CAP prunes by incorporating communication cost into routing-destination removal decisions, reducing inter-device communication while limiting accuracy degradation to achieve the target speedup.\"}]",1784183874,30,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"communication-aware-placement-and-pruning-for-efficient-mixture-of-experts-inference","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/communication-aware-placement-and-pruning-for-efficient-mixture-of-experts-inference/82909/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does CAP address in MoE inference?","Question",{"text":75,"@type":76},"CAP targets sub-optimal MoE inference caused by expert placement and pruning decisions that ignore their effect on communication volume, which becomes a dominant bottleneck at large expert counts.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does CAP determine expert placement?",{"text":80,"@type":76},"CAP uses co-activation-driven placement to group frequently co-activated experts, then generates candidate placements by adjusting the communication–computation trade-off to explore different load and communication balances.",{"name":82,"@type":73,"acceptedAnswer":83},"What is communication-aware expert pruning in CAP?",{"text":84,"@type":76},"CAP prunes by incorporating communication cost into routing-destination removal decisions, reducing inter-device communication while limiting accuracy degradation to achieve the target speedup.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":29,"slug":121},"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]