[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85329-en":3,"doc-seo-85329-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85329,1099514068365,"Aurelia","https://ap-avatar.wpscdn.com/avatar/10000253d8d9f28188e?_k=1776742907772140068",8,"Research & Report","Unlocking Every Expert in Domain-Specific Training","Mixture-of-Experts (MoE) models increase model capacity without proportional compute cost and have become central to frontier LLMs, yet domain-specific post-training inherits an expert pool shaped by mixed-domain pre-training. A large subset of experts contributes little to the target domain, and standard SFT updates experts without reassigning expert slots. UMoE realigns the expert pool via pruning low domain-aligned saliency experts, perturbation-based regrowth, and then applying standard SFT, preserving parameter count, expert count, and inference cost while improving results across multiple architectures, domains, and benchmarks.","arXiv :2607 . 1 1444v 1 [ cs .CL] 13 Jul 2026  \nUnlocking Every Expert in Domain-Specific Training  \nXuefeng Li 1,2,3 Pengfei Liu†1,2,3  \n1 SII 2 SJTU 3 GAIR  \nAbstract  \nMixture-of-Experts (MoE) models scale capacity without proportional compute cost and have become a key architecture for frontier large language models (LLMs) . Yet domain-specific post-training inherits an expert pool shaped by mixed-domain pre-training: a substantial subset of experts contributes little on the target domain, and standard supervised fine-tuning (SFT) leaves the composition of this pool unchanged. We propose a simple, budget-preserving pipeline that realigns the expert pool to the target domain before fine-tuning. Given a target domain, we (1) prune the experts with lowest domain-aligned saliency,(2) regrow the expert pool to its original size through perturbation-based expert expansion, and (3) apply standard SFT. The resulting model preserves the original expert count, parameter count, and inference cost. With a single frozen recipe and no per-domain hyperparameter tuning, UMoE consistently improves over Direct SFT across two MoE architectures (Qwen3-30B-A3B and Qwen3.5-35B-A3B), five domains (math, code, science, tool-use, and agentic coding), and 12 benchmarks. Representative improvements are 3.4 points in math average accuracy, 6.0 points on SWE-bench Verified. On a strong in-house math corpus, Direct SFT already surpasses Qwen3-30B-A3B-Thinking (82.81 vs. 81.06), yet UMoE further raises the average to 84.17, an additional 1.36 points, demonstrating robustness to a substantially stronger SFT regime. Data-scaling experiments further show that the gain persists as training data grows. Analysis reveals that the direct-SFT model allocates substantial routed-expert compute to a low-saliency subset that can be removed post hoc with little average degradation; UMoE turns this redundant capacity into useful domain capacity and achieves lower training loss, with gains spanning all difficulty levels in downstream evaluation.  \nFigure 1: Overview of UMoE. Direct SFT fine-tunes the pretrained MoE as is. UMoE first prunes the least domain-salient experts and regrows the pool from the retained experts, restoring the original size and inference cost before SFT. The same domain data D is used for calibration and training.  \n†Corresponding author.  \n1 Introduction  \nFrontier large language models such as Qwen3.5 (Qwen Team, 2026a), DeepSeek-V4 (DeepSeek-AI, 2026), and Kimi-K2 (Team, 2025) increasingly adopt Mixture-of-Experts (MoE) architectures (Shazeer et al., 2017) to expand model capacity while keeping inference efficient. Building on these models, domain-specific post-training serves two important purposes. It can directly push capability on the hardest problems within a target domain (Shao et al., 2025 ; Team, 2026); it can also produce specialist models whose capabilities are later consolidated into a generalist through on-policy distillation (Lu and Thinking Machines Lab, 2025 ; Ma et al., 2026) or rejection-sampling fine-tuning (Yang et al., 2025 ; DeepSeek-AI, 2025 ; LLM-Core Xiaomi, 2026 ; Huang et al., 2026 ; DeepSeek-AI, 2026) . Both settings require adapting a general-purpose MoE to a specific target domain.  \nThe expert pool inherited from mixed-domain pre-training, however, is not homogeneous. Sparse routing over diverse data encourages experts to develop different functional specializations (Dai et al., 2024 ; Muennighoff et al., 2024), making their relevance uneven under a single target distribution. Consequently, a low-saliency subset can continue to receive non-trivial routing on domain data. Standard SFT updates these inherited experts in place but does not explicitly reallocate expert slots, leaving part of the fixed capacity poorly aligned with the target domain.  \nWe therefore ask: can the expert pool be reorganized into a more domain-aligned initialization before SFT, so that more expert slots become useful for learning target-dom","cbCaikzi4lWdIwio","https://ap.wps.com/l/cbCaikzi4lWdIwio","pdf",658528,2,1,13,"English","en",105,"# Abstract\n# Introduction\n# Related Work","[{\"question\":\"What problem does UMoE address in domain-specific MoE post-training?\",\"answer\":\"It addresses the mismatch between a target domain and the inherited expert pool from mixed-domain pre-training, where a low-saliency subset of experts still receives routing during domain training and remains poorly aligned under standard SFT.\"},{\"question\":\"How does UMoE realign experts before fine-tuning?\",\"answer\":\"UMoE calibrates on target-domain data to prune the lowest domain-aligned saliency experts, then regrows the pool using perturbation-based expert expansion to restore the original expert count and inference cost, and finally runs standard SFT.\"},{\"question\":\"Does UMoE keep the model’s computational budget unchanged?\",\"answer\":\"Yes. UMoE preserves the original expert count, parameter count, and inference cost by pruning and regrowing back to the same pool size before applying SFT.\"}]",1784202528,33,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"unlocking-every-expert-in-domain-specific-training","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/unlocking-every-expert-in-domain-specific-training/85329/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-23","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does UMoE address in domain-specific MoE post-training?","Question",{"text":75,"@type":76},"It addresses the mismatch between a target domain and the inherited expert pool from mixed-domain pre-training, where a low-saliency subset of experts still receives routing during domain training and remains poorly aligned under standard SFT.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does UMoE realign experts before fine-tuning?",{"text":80,"@type":76},"UMoE calibrates on target-domain data to prune the lowest domain-aligned saliency experts, then regrows the pool using perturbation-based expert expansion to restore the original expert count and inference cost, and finally runs standard SFT.",{"name":82,"@type":73,"acceptedAnswer":83},"Does UMoE keep the model’s computational budget unchanged?",{"text":84,"@type":76},"Yes. UMoE preserves the original expert count, parameter count, and inference cost by pruning and regrowing back to the same pool size before applying SFT.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]