[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-seo-203803-105":3,"detail-sidebar-cat-0-en-105":81,"doc-detail-203803-en":131},{"code":4,"msg":5,"data":6},0,"ok",{"site_id":7,"language":8,"slug":9,"title":10,"keywords":11,"description":12,"schema_data":13,"social_meta":74,"head_meta":76,"extra_data":78,"updated_unix":80},105,"en","sere-similarity-based-expert-re-routing-for-efficient-batch-decoding-in-moe-models","SERE: Similarity-based Expert Re-routing for Efficient Batch Decoding in MoE Models","","Mixture-of-Experts (MoE) architectures use sparse activation to improve training and inference efficiency compared with dense LLMs, yet production serving often relies on batch inference. Batching can activate too many experts, increasing overhead and slowing the memory-bound decoding stage. SERE introduces similarity-based expert re-routing that token-awarely redirects tokens from secondary experts to the most similar primary counterparts. It also identifies and preserves critical experts via similarity patterns, avoiding capability loss, while using an efficient custom CUDA kernel for plug-and-play deployment in vLLM.",{"@graph":14,"@context":73},[15,34,56],{"@type":16,"itemListElement":17},"BreadcrumbList",[18,23,27,31],{"item":19,"name":20,"@type":21,"position":22},"https://docshare.wps.com","Home","ListItem",1,{"item":24,"name":25,"@type":21,"position":26},"https://docshare.wps.com/document/","Document",2,{"item":28,"name":29,"@type":21,"position":30},"https://docshare.wps.com/document/research-report/","Research & Report",3,{"item":32,"name":10,"@type":21,"position":33},"https://docshare.wps.com/document/sere-similarity-based-expert-re-routing-for-efficient-batch-decoding-in-moe-models/203803/",4,{"url":32,"name":10,"@type":35,"image":36,"author":41,"headline":10,"publisher":44,"fileFormat":47,"inLanguage":8,"description":12,"dateModified":48,"datePublished":49,"encodingFormat":47,"isAccessibleForFree":50,"interactionStatistic":51},"DigitalDocument",{"url":37,"@type":38,"width":39,"height":40},"https://docshare.wps.com/thumbnails/sere-similarity-based-expert-re-routing-for-efficient-batch-decoding-in-moe-models/203803.png","ImageObject",300,407,{"name":42,"@type":43},"Emma Mercer","Person",{"url":19,"name":45,"@type":46},"DocShare","Organization","application/pdf","2026-10-08","2026-09-04",true,{"@type":52,"interactionType":53,"userInteractionCount":55},"InteractionCounter",{"@type":54},"ViewAction",11,{"@type":57,"mainEntity":58},"FAQPage",[59,65,69],{"name":60,"@type":61,"acceptedAnswer":62},"What problem does SERE target in MoE production serving?","Question",{"text":63,"@type":64},"SERE targets the mismatch between selective expert activation and batched inference, where batches can activate excessive experts and increase memory/communication overhead during decoding.","Answer",{"name":66,"@type":61,"acceptedAnswer":67},"How does SERE reduce the number of active experts?",{"text":68,"@type":64},"SERE re-routes tokens dynamically in an input-aware way by redirecting tokens from secondary experts to their most similar primary counterparts based on similarity patterns.",{"name":70,"@type":61,"acceptedAnswer":71},"Does SERE rely on static expert pruning or merging?",{"text":72,"@type":64},"No. SERE avoids static expert pruning or merging and instead enables dynamic expert skipping driven by batch-level expert redundancy.","https://schema.org",{"og:url":32,"og:type":75,"og:title":10,"og:site_name":45,"og:description":12},"article",{"robots":77,"canonical":32},"index,follow",{"doc_id":79,"site_id":7},203803,1788563991,{"code":4,"msg":82,"data":83},"success",[84,88,92,96,101,106,111,115,120,123,127],{"id":22,"doc_module":4,"doc_module_name":25,"category_name":85,"show_sort_weight":86,"slug":87},"Story & Novel",90,"story-novel",{"id":26,"doc_module":4,"doc_module_name":25,"category_name":89,"show_sort_weight":90,"slug":91},"Literature",80,"literature",{"id":33,"doc_module":4,"doc_module_name":25,"category_name":93,"show_sort_weight":94,"slug":95},"Exam",70,"exam",{"id":97,"doc_module":4,"doc_module_name":25,"category_name":98,"show_sort_weight":99,"slug":100},5,"Comic",60,"comic",{"id":102,"doc_module":4,"doc_module_name":25,"category_name":103,"show_sort_weight":104,"slug":105},6,"Technology",50,"technology",{"id":107,"doc_module":4,"doc_module_name":25,"category_name":108,"show_sort_weight":109,"slug":110},7,"Healthcare",40,"healthcare",{"id":112,"doc_module":4,"doc_module_name":25,"category_name":29,"show_sort_weight":113,"slug":114},8,30,"research-report",{"id":116,"doc_module":4,"doc_module_name":25,"category_name":117,"show_sort_weight":118,"slug":119},9,"Religion & Spirituality",20,"religion-spirituality",{"id":118,"doc_module":4,"doc_module_name":25,"category_name":121,"show_sort_weight":118,"slug":122},"World Cup","world-cup",{"id":124,"doc_module":4,"doc_module_name":25,"category_name":125,"show_sort_weight":124,"slug":126},10,"Lifestyle","lifestyle",{"id":128,"doc_module":4,"doc_module_name":25,"category_name":129,"show_sort_weight":97,"slug":130},19,"General","general",{"code":4,"msg":82,"data":132},{"doc_id":79,"user_id":133,"nickname":42,"user_avatar":134,"doc_module":4,"category_id":112,"category_name":29,"doc_title":10,"doc_description":12,"doc_content":135,"file_id":136,"file_url":137,"file_type":138,"file_size":139,"view_count":55,"is_deleted":4,"is_public":22,"is_downloadable":22,"audit_status":22,"page_count":140,"language":141,"language_code":8,"site_id":7,"html_lang":8,"table_of_contents":142,"faqs":143,"seo_title":144,"seo_description":12,"update_tm":80,"read_time":145},962084925502,"https://ap-avatar.wpscdn.com/davatar_6f874abed73319feea01a86fa6f0fab8","arXiv :2602 .076 16v 1 [ cs .LG] 7 Feb 2026  \nSERE: SIMILARITY-BASED EXPERT RE-ROUTING FOR EFFICIENT BATCH DECODING IN MOE MODELS  \nJuntong Wu1,2,* , Jialiang Cheng1,*,†, B, Fuyu Lv1 , Ou Dan1 , Li Yuan2, B  \n1 Taobao & Tmall Group of Alibaba  \n2 Shenzhen Graduate School, Peking University  \nCorrespondence: [jichen.cjl@alibaba-inc.com](jichen.cjl@alibaba-inc.com), [yuanli-ece@pku.edu.cn](yuanli-ece@pku.edu.cn)  \nABSTRACT  \nMixture-of-Experts (MoE) architectures employ sparse activation to deliver faster training and inference with higher accuracy than dense LLMs. However, in production serving, MoE models require batch inference to optimize hardware efficiency, which may cause excessive expert activation and thus slow the memorybound decoding stage. To address the fundamental tension between batch decoding and expert sparsity, we present SERE, a Similarity-based Expert Re-routing method for Efficient batch decoding in MoE models. SERE dynamically reduces the number of active experts in an input-aware manner by re-routing tokens from secondary experts to their most similar primary counterparts. It also leverages similarity patterns to identify and preserve critical experts, thereby preventing capability loss. Notably, SERE avoids static expert pruning or merging, instead enabling dynamic expert skipping based on batch-level expert redundancy. Additionally, we provide an efficient custom CUDA kernel for SERE, enabling plug-and-play use in vLLM with only a single-line code change.1 Extensive experiments on various complex reasoning benchmarks demonstrate that SERE achieves up to 2.0× speedup with minimal quality loss, providing a practical solution for cost-efficient and latency-sensitive large-scale MoE deployment.  \n1 INTRODUCTION  \nLarge Language Models (LLMs) have shown remarkable performance across various applications. Recently, the Mixture-of-Experts (MoE) paradigm has emerged as a leading framework for scaling LLMs (Yang et al., 2025a; Liu et al., 2024b; Touvron et al., 2023; Jiang et al., 2024) . Unlike dense LLMs that activate the entire feed-forward network (FFN) for every token, an MoE layer consists of multiple lightweight FFN experts, where a learnable router assigns each token to a small subset. By maintaining low per-token computation, sparse activation enables the model to incorporate numerous specialized experts, scaling its capacity while preserving training and inference efficiency.  \nDespite the theoretical efficiency of MoE architectures, their practical gains are often limited by a mismatch between selective activation and batched inference (Kwon et al., 2023; Agrawal et al., 2024; Gupta et al., 2024) . In real-world services, multiple user requests are batched to improve hardware utilization (Kwon et al., 2023) . However, tokens within a batch often require different experts, leading to a total number of activated experts far above the per-token budget (Agrawal et al., 2024; Yun et al., 2024) . As depicted in Figure 1, even with strict limits (e.g., 8 out  \nAverage Activated Experts  \n* Equal contribution † Project Lead B Corresponding author  \n1 Code implementation of SERE can be found in [https://github.com/JL-Cheng/SERE](https://github.com/JL-Cheng/SERE).  \nMATH  \n72.2  \n53.4  \nBoolQ 90.2  \n64.5  \n54.2  \n3.3 6.9 16.9  \n10.4  \n10.1  \n54.8 53.2  \n79.0 78.4  \nMATH401 64.0 MBPP  \nBBH  \n76.7  \n65.4  \nGSM8K 89.2  \nCMMLU 84.8  \n87.2 HumanEval  \n(a) Accuracy Maintained  \nOrigin LYNX SERE Origin LYNX SERE Origin LYNX SERE  \n(b) Time Per Output Token (ms)  \nFigure 2: Visualizations of SERE’s Performance. (a) Across all tasks, SERE (K=2) exhibits negligible performance loss, while SERE (K=1) still outperforms all baselines. (b) SERE significantly reduces batch decoding time, achieving up to 2× acceleration.  \nof 128 in Qwen3-30B-A3B (Yang et al., 2025a)), a moderately diverse batch can still activate a majority of the experts simultaneously. Moreover, the training-time load-balancing objectives further increase th","cbCaifIuwWShw3gq","https://ap.wps.com/l/cbCaifIuwWShw3gq","pdf",8969285,27,"English","# Abstract\n# Introduction\n## Background: MoE sparse activation and routing\n## Problem: batch decoding vs expert sparsity\n## Related work: static compression and dynamic skipping\n## Proposed method: SERE","[{\"question\":\"What problem does SERE target in MoE production serving?\",\"answer\":\"SERE targets the mismatch between selective expert activation and batched inference, where batches can activate excessive experts and increase memory/communication overhead during decoding.\"},{\"question\":\"How does SERE reduce the number of active experts?\",\"answer\":\"SERE re-routes tokens dynamically in an input-aware way by redirecting tokens from secondary experts to their most similar primary counterparts based on similarity patterns.\"},{\"question\":\"Does SERE rely on static expert pruning or merging?\",\"answer\":\"No. SERE avoids static expert pruning or merging and instead enables dynamic expert skipping driven by batch-level expert redundancy.\"}]","SERE: Similarity-based Expert Re-routing for Efficient Batch Decoding in MoE Models | PDF",68]