[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82134-en":3,"doc-seo-82134-105":29,"detail-sidebar-cat-0-en-105":83},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},82134,1099514067438,"River Wang","https://ap-avatar.wpscdn.com/avatar/100002539ee87300030?x-image-process=image/resize,m_fixed,w_180,h_180&k=1780474512215547542",8,"Research & Report","Sensitivity Aware Thresholding and Token Routing for Activation Sparsification in Large Language Models","Efficient inference in Large Language Models (LLMs) depends on reducing computation without harming quality. The work studies multilayer perceptron (MLP) activation sparsification combined with token-level conditional routing. It proposes SATS (Sensitivity-Aware Thresholding for Sparsity), which calibrates layerwise gate thresholds using a local MLP output sensitivity proxy rather than activation percentiles. A lightweight token routing framework selects dense vs sparse paths per token. Experiments on open-weight LLMs show improved quality at matched sparsity and a better quality-throughput trade-off than static baselines.","Sensitivity-Aware Thresholding and Token Routing for Activation Sparsification in Large Language Models  \nBishmoy Paul Santa Clara University [bishmoypaul. contact@gmail. com](bishmoypaul. contact@gmail. com)  \nYoungmin Yi Sogang University [ymyi@sogang. ac. kr](ymyi@sogang. ac. kr)  \nHoeseok Yang Santa Clara University [hoeseok. yang@scu. edu](hoeseok. yang@scu. edu)  \narXiv :2607 .0899 1v 1 [ cs .LG] 9 Jul 2026  \nAbstract  \nEfficient inference in Large Language Models (LLMs) requires deciding where computation can be reduced while preserving model quality. We study this problem through multilayer perceptron (MLP) activation sparsification and token-level conditional routing. We first propose Sensitivity-Aware Thresholding for Sparsity (SATS), a threshold calibration method to choose layerwise gate thresholds using a local MLP output sensitivity proxy rather than calibrating thresholds directly from activation percentiles. While SATS retains the existing mechanism of sparsifying MLP activations by thresholding gate activations, it replaces percentile-based calibration with a sensitivity-aware selection rule. We then introduce a lightweight token routing framework that dynamically selects between a base path and a modified path on a per-token basis, rather than applying the modified computation uniformly to all tokens. We evaluate both methods on multiple recent open-weight LLMs. Our results show that SATS improves over the threshold-based sparsification baseline at matched actual sparsity and that token routing yields a more favorable quality-throughput trade-off than static activation modification baselines. Overall, our results suggest that improved threshold calibration and token routing can improve the quality-throughput trade-off in LLMs.  \n1 Introduction  \nLarge Language Models (LLMs) are widely used across various tasks, but their inference cost remains a major bottleneck. A significant portion of this cost comes from the multilayer perceptron (MLP) blocks inside each transformer layer. Therefore, optimizing MLP blocks can speed up on-device LLM inference. Recent works (Lee et al. , 2024; Zhang et al. , 2024) have shown that these activations contain significant sparsity that can be exploited by thresholding MLP gate activations, which can speed up LLM inference without modifying the overall transformer architecture. However, deciding which activations to remove and when to apply the modified computation remains a challenge.  \nA common strategy to address this challenge is to calibrate one activation threshold per layer from activation statistics from a calibration dataset. CATS (Lee et al. , 2024) uses a percentile-based approach to choose these layerwise thresholds. But percentile-based calibration focuses primarily on how many activations are removed, instead of examining how much damage a given threshold causes to the resultant MLP output. Two adjacent thresholds in the activation layer can have drastically different amounts of information lost in the final output. Even more broadly, activation modification is applied uniformly across all tokens for a generation prompt, even though the impact of the modification is not necessarily uniform across tokens.  \nIn this work, we study sparsification of the MLP layer in LLMs at two levels. First, we propose Sensitivity-aware Thresholding for Sparsity (SATS), a threshold calibration method that replaces percentile-based threshold selection with a sensitivity-aware method based on layerwise  \nMLP output distortion. SATS chooses thresholds to satisfy a local layer sensitivity budget, and the layerwise final operating points are chosen to satisfy a global target sparsity. Secondly, we introduce a lightweight token routing method that decides whether to use the dense MLP path or the sparse path (that uses the activation thresholding) for each token. The routing method only requires the token identity, which allows it to be simplified into a look-up table with negligible effect ","cbCaiihR8WUoPMQ8","https://ap.wps.com/l/cbCaiihR8WUoPMQ8","pdf",1290391,1,12,"English","en",105,"# Introduction\n# Background","[{\"question\":\"What evaluation results are reported for the proposed methods?\",\"answer\":\"Experiments on models such as Llama 3.1 8B and Qwen 3 8B show SATS improves quality over percentile-based sparsification at matched realized sparsity, and token routing improves the quality-throughput trade-off versus static sparse or dense execution.\"}]",1784178384,30,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":78,"head_meta":80,"extra_data":82,"updated_unix":27},"sensitivity-aware-thresholding-and-token-routing-for-activation-sparsification-in-large-language-models","",{"@graph":35,"@context":77},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/sensitivity-aware-thresholding-and-token-routing-for-activation-sparsification-in-large-language-models/82134/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71],{"name":72,"@type":73,"acceptedAnswer":74},"What evaluation results are reported for the proposed methods?","Question",{"text":75,"@type":76},"Experiments on models such as Llama 3.1 8B and Qwen 3 8B show SATS improves quality over percentile-based sparsification at matched realized sparsity, and token routing improves the quality-throughput trade-off versus static sparse or dense execution.","Answer","https://schema.org",{"og:url":51,"og:type":79,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":81,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":84},[85,89,93,97,102,107,112,114,119,122,126],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":86,"show_sort_weight":87,"slug":88},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":90,"show_sort_weight":91,"slug":92},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Exam",70,"exam",{"id":98,"doc_module":4,"doc_module_name":45,"category_name":99,"show_sort_weight":100,"slug":101},5,"Comic",60,"comic",{"id":103,"doc_module":4,"doc_module_name":45,"category_name":104,"show_sort_weight":105,"slug":106},6,"Technology",50,"technology",{"id":108,"doc_module":4,"doc_module_name":45,"category_name":109,"show_sort_weight":110,"slug":111},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":28,"slug":113},"research-report",{"id":115,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":117,"slug":118},9,"Religion & Spirituality",20,"religion-spirituality",{"id":117,"doc_module":4,"doc_module_name":45,"category_name":120,"show_sort_weight":117,"slug":121},"World Cup","world-cup",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":123,"slug":125},10,"Lifestyle","lifestyle",{"id":127,"doc_module":4,"doc_module_name":45,"category_name":128,"show_sort_weight":98,"slug":129},19,"General","general"]