[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83781-en":3,"doc-seo-83781-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":11,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},83781,4398048950312,"Violet","https://ap-avatar.wpscdn.com/avatar/400002538284de19e3c?_k=1778320343897328908",8,"Research & Report","BrownoutMoE Structure-Aware Expert Grouping for Efficient and Accurate LLM Web-based Services","BrownoutMoE addresses efficient and accurate inference for Mixture-of-Experts (MoE) LLMs deployed in Web-facing services, where bursty traffic requires both responsiveness and quality. The core challenge is structurally imbalanced expert access: a few hot experts dominate routed tokens while many cold experts remain unused, wasting GPU parallelism. BrownoutMoE reorganizes experts into united groups via layer-wise reinforcement learning, then applies grouping-consistent distillation for deployable models. Experiments show up to 71.4% less accuracy degradation and up to 2.24× higher throughput than baselines.","arXiv :2607 .04 164v 1 [ cs .DC] 5 Jul 2026  \nBrownoutMoE: Structure-Aware Expert Grouping for Efficient and Accurate LLMWeb-based Services  \nYi Ding 1 ,2 , Minxian Xu 1 ,(􀀌), Zhengxin Fang3 , Kejiang Ye 1 , and Chengzhong  \nXu4  \n1 Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences,  \nShenzhen, China  \n{yi.ding2,[mx.xu}@siat.ac.cn](mx.xu}@siat.ac.cn) , [zhengxin.fang@ecs.vuw.ac.nz](zhengxin.fang@ecs.vuw.ac.nz) ,  \n[kj.ye@siat.ac.cn](kj.ye@siat.ac.cn) , [czxu@um.edu.mo](czxu@um.edu.mo)  \n2 University of Chinese Academy of Sciences, Beijing, China  \n3 Victoria University of Wellington, Wellington, New Zealand  \n4 State Key Lab of IoTSC, University of Macau, Macau, China  \nAbstract. Mixture-of-Experts (MoE) large language models (LLMs) are increasingly deployed in Web-facing services, where inference must be both accurate and responsive under bursty demand. Although MoE models improve parameter efficiency through sparse expert activation, efficient MoE inference remains challenging in practice. A major reason is the highly imbalanced expert access pattern during inference: a few hot experts process most routed tokens, while many cold experts are rarely activated, leaving GPU parallelism underutilized. Existing systems mainly optimize runtime execution, such as scheduling, communication overlap, and kernel fusion, but usually preserve the original expert organization and therefore do not address the structural inefficiency caused by fragmented expert usage. In this paper, we present BrownoutMoE, a structure-aware optimization framework for efficient and accurate MoE inference services. Inspired by the brownout paradigm in service computing, BrownoutMoE reorganizes experts into groups to improve utilization and system efficiency while maintaining service quality. Specifically, we formulate layer-wise expert grouping as a learning problem and employ reinforcement learning to discover grouping strategies that minimize accuracy degradation. We further introduce a grouping-consistent distillation process to produce deployable models that are compatible with standard inference pipelines. Experimental results demonstrate that BrownoutMoE reduces accuracy degradation by up to 71 .4% and improves throughput by up to 2.24× over baselines.  \nKeywords: LLM Web-Based Applications · MoE Inference Serving · Expert Grouping · GRPO · Knowledge Distillation  \n1 Introduction  \nLarge language models (LLMs) are increasingly used in Web-based services such as search, recommendation, question answering, fact checking, and interactive  \n2 Y. Ding et al.  \nWeb agents, while cloud-native and distributed systems have become central to scalable LLM deployment [14] . In these LLM Web-based applications, inference is no longer an offline batch task but part of the online request path. User-facing services must handle bursty arrivals, heterogeneous prompt and response lengths, and session-level interactions while maintaining response-time and answer-quality expectations.  \nThese Web service characteristics create a demanding serving problem. The runtime must jointly balance throughput, tail latency, and accuracy under dynamic workloads. Mixture-of-Experts (MoE) architectures [8,10] are attractive because they scale model capacity through sparse expert activation. However, Web-facing workloads can amplify skewed expert access: a few experts receive most routed tokens, while many others are rarely activated. This imbalance creates hot-expert bottlenecks and poor GPU utilization for cold experts.  \nExisting MoE serving systems mainly improve execution efficiency through expert parallelism [3], scheduling [15], and fused kernels [4] . However, these methods usually preserve the original expert organization, leaving the structural inefficiency of fragmented expert usage unchanged. BrownoutServe [7] introduced united experts to reduce expert access frequency, but its grouping strategy is still based on fixed heuristics.  \nIn this work, we argue ","cbCaitvijJPgU1VE","https://ap.wps.com/l/cbCaitvijJPgU1VE","pdf",4347933,1,15,"English","en",105,"# Abstract\n# Introduction\n## Challenges in MoE Web Serving\n## Proposed BrownoutMoE Framework\n## Contributions","[{\"question\":\"Why is MoE inference inefficient in Web-facing services?\",\"answer\":\"Inference often routes most tokens to a small set of hot experts, while many cold experts are rarely activated. This creates hot-expert bottlenecks and underutilizes GPU parallelism.\"},{\"question\":\"How does BrownoutMoE optimize expert grouping?\",\"answer\":\"BrownoutMoE formulates layer-wise expert grouping as a discrete policy optimization problem and uses GRPO to learn grouping strategies that reduce accuracy degradation after distillation.\"},{\"question\":\"What role does grouping-consistent distillation play?\",\"answer\":\"After GRPO finds grouping maps, grouping-consistent united-expert distillation produces deployable models that remain compatible with standard inference pipelines while preserving expert behavior.\"}]",1784190363,38,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"brownoutmoe-structure-aware-expert-grouping-for-efficient-and-accurate-llm-web-based-services","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/brownoutmoe-structure-aware-expert-grouping-for-efficient-and-accurate-llm-web-based-services/83781/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":11},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why is MoE inference inefficient in Web-facing services?","Question",{"text":75,"@type":76},"Inference often routes most tokens to a small set of hot experts, while many cold experts are rarely activated. This creates hot-expert bottlenecks and underutilizes GPU parallelism.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does BrownoutMoE optimize expert grouping?",{"text":80,"@type":76},"BrownoutMoE formulates layer-wise expert grouping as a discrete policy optimization problem and uses GRPO to learn grouping strategies that reduce accuracy degradation after distillation.",{"name":82,"@type":73,"acceptedAnswer":83},"What role does grouping-consistent distillation play?",{"text":84,"@type":76},"After GRPO finds grouping maps, grouping-consistent united-expert distillation produces deployable models that remain compatible with standard inference pipelines while preserving expert behavior.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]