[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86273-en":3,"doc-seo-86273-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":11,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},86273,1099514068365,"Aurelia","https://ap-avatar.wpscdn.com/avatar/10000253d8d9f28188e?_k=1776742907772140068",8,"Research & Report","HCRMap: Pressure-Aware Hot-Expert Residency Mapping for 3.5D MoE Chiplet Inference","Mixture-of-Experts (MoE) LLM inference routes tokens to only a small subset of experts, yet token-dependent routing creates persistent hotness skew where a few hot experts dominate traffic while others remain lightly used. In 3.5D multi-chiplet systems, this skew simultaneously causes compute imbalance and escalates pressure on communication, memory bandwidth, I/O, and execution queues. HCRMap addresses this by dynamically managing expert replica residency across memory tiers. It promotes, retains, demotes, or evicts experts and maps routed token groups to resident replicas to mitigate bottlenecks.","arXiv :2607 . 1 1586v 1 [ cs .AI] 13 Jul 2026  \nHCRMap: Pressure-Aware Hot-Expert Residency Mapping for 3.5D MoE Chiplet Inference  \nYongqin Zhanga,∗  \na Nanjing Vocational College of Information Technology, 99 Wenlan Road, Xianlin College Community, Nanjing, Jiangsu, 210023, P.R.C.  \nARTICLE INFO  \nKeywords:  \nMixture-of-Experts 3.5D chiplets Runtime mapping Hot expert residency Hierarchical memory  \nAB STRACT  \nMixture-of-Experts (MoE) large language models (LLM) activate only a small number of experts during inference, but token routing introduces persistent expert hotness skew: a small set of hot experts continuously receives most tokens, while the remaining experts are lightly loaded. On 3.5D multi-chiplet systems, this skew not only causes compute imbalance but also amplifies pressure on communication, memory bandwidth, I/O, and execution queues. Therefore, the core problem is not simply to reduce token movement, but to dynamically place and reuse hot expert replicas across different memory tiers.  \nThis paper proposes HCRMap, a hot expert residency mapping framework for pressure-aware expert replica management in 3.5D MoE inference. Based on expert hotness, weight loading cost, migration overhead, and runtime resource pressure, HCRMap dynamically determines which experts should be promoted, retained, demoted, or evicted. It then maps routed token groups to suitable resident replicas, thereby jointly mitigating communication, memory, and queue bottlenecks.  \nExperimental results show that HCRMap reduces end-to-end latency by 43.6% and 43.0% over Hydra in the prefill and decode stages, respectively; by 34.5% and 33.1% over MoEntwine; and by 46.7% and 46 .0% over PIMoE.  \n1. Introduction  \nMixture-of-Experts large language models (MoELLMs) scale capacity by activating only a few experts per token through Top-􀁫 gating [1–4], but token-dependent routing produces persistent expert-popularity skew: a small set of hot experts repeatedly receives most tokens across serving windows while others stay lightly used [5–8] . As shown in Figure 1, this skew is not only a compute-load problem. It reshapes token dispatch and gather, expertweight access, memory-bank pressure, and shared-I/O pressure. The routed expert-FFN bottleneck is therefore a cross-resource imbalance over compute, memory, and interconnect, affecting full-model latency and energy in both prefill and decode execution.  \nOn 3.5D multi-chiplet systems with hierarchical memory, this skew surfaces as repeated expert-weight streaming [9–11] . Expert weights are too large to keep all active experts in the nearest stacked-SRAM tier; in our evaluation, each benchmark uses the routed-expert footprint computed from its own hidden and intermediate dimensions, while the near-tier budget of each chiplet group is only 64 MB. Therefore, finite near-tier capacity can hold only a small number of hot experts, while routing profiles may require more hot or warm resident experts in the same layer. Warm experts that overflow the near tier must repeatedly stream from groupshared DRAM through shared I/O and D2D/NoP links. This transfer recurs across windows and continuously occupies shared interfaces and package links. Because 3.5D integration co-packages stacked SRAM, local HBM, and groupshared DRAM into a finer-grained hierarchy, expert weights should not be classified simply as resident or non-resident;  \n∗Corresponding author.  \nEmail address: [zyq1350335@126.com](zyq1350335@126.com) (Y. Zhang)  \nFigure 1: MoE FFN serving pipeline for selected experts.  \nthey can be placed at different residency levels according to runtime hotness, streaming cost, and resource pressure.  \nExisting MoE systems optimize expert execution along separate axes: placement and routing that use popularity, affinity, or co-activation to cut dispatch latency; replication or shadow replicas that relieve hot-expert queueing; and heterogeneous execution or memory-aware offloading driven by compute and memory behavior ","cbCaincMFLLKty4t","https://ap.wps.com/l/cbCaincMFLLKty4t","pdf",1699630,1,15,"English","en",105,"# Introduction\n## Hotness skew in MoE inference\n## Limitations of existing systems\n## Residency placement as a cross-resource tradeoff","[{\"question\":\"What problem does HCRMap address in 3.5D MoE chiplet inference?\",\"answer\":\"HCRMap targets persistent hot-expert skew that creates cross-resource imbalance across compute, memory, communication, I/O, and execution queues in 3.5D multi-chiplet systems.\"},{\"question\":\"How does HCRMap decide which expert replicas to keep or move?\",\"answer\":\"HCRMap dynamically promotes, retains, demotes, or evicts experts using hotness, weight loading cost, migration overhead, and runtime resource pressure.\"},{\"question\":\"What performance improvements does the paper report for HCRMap?\",\"answer\":\"Experiments show HCRMap reduces end-to-end latency by 43.6% and 43.0% over Hydra in prefill and decode, respectively, and also improves versus MoEntwine and PIMoE.\"}]",1784209969,38,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"hcrmap-pressure-aware-hot-expert-residency-mapping-for-35d-moe-chiplet-inference","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/hcrmap-pressure-aware-hot-expert-residency-mapping-for-35d-moe-chiplet-inference/86273/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":11},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does HCRMap address in 3.5D MoE chiplet inference?","Question",{"text":75,"@type":76},"HCRMap targets persistent hot-expert skew that creates cross-resource imbalance across compute, memory, communication, I/O, and execution queues in 3.5D multi-chiplet systems.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does HCRMap decide which expert replicas to keep or move?",{"text":80,"@type":76},"HCRMap dynamically promotes, retains, demotes, or evicts experts using hotness, weight loading cost, migration overhead, and runtime resource pressure.",{"name":82,"@type":73,"acceptedAnswer":83},"What performance improvements does the paper report for HCRMap?",{"text":84,"@type":76},"Experiments show HCRMap reduces end-to-end latency by 43.6% and 43.0% over Hydra in prefill and decode, respectively, and also improves versus MoEntwine and PIMoE.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]