[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-122509-en":3,"doc-seo-122509-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},122509,687197207919,"Theodora","https://ap-avatar.wpscdn.com/avatar/a000253d6f5f7c60be?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779446848396160552",8,"Research & Report","Pushing the Limits of Sparsity - A Bag of Tricks for Extreme Pruning","Pruning deep neural networks reduces model size while largely preserving dense-model performance, enabling deployment on memory- and power-constrained devices. Prior sparse learning reaches strong results up to moderate sparsity (about 95%–98%), but accuracy degrades rapidly at extreme sparsity due to instability such as fragile gradient flow. This work studies performance beyond commonly explored sparsities and introduces Extreme Adaptive Sparse Training (EAST). EAST achieves stable accuracy at 99.90%, 99.95%, and 99.99% on ResNet architectures via Dynamic ReLU phasing, parameter weight sharing, and cyclic sparsity scheduling, evaluated on CIFAR-10, CIFAR-100, and ImageNet.","Pushing the Limits of Sparsity: A Bag of Tricks for Extreme Pruning  \nAndy Li  \nDepartment of Computing Science University of Aberdeen, UK  \nAiden Durrant  \nDepartment of Computing Science University of Aberdeen, UK  \nMilan Markovic  \nDepartment of Computing Science & Interdisciplinary Institute University of Aberdeen, UK  \nTianjin Huang  \nDepartment of Computer Science University of Exeter, UK  \nSouvik Kundu  \nIntel Labs, USA  \nTianlong Chen  \nDepartment of Computer Science  \nUniversity of North Carolina at Chapel Hill, US  \nLu Yin  \nSchool of Computer Science and Electronic Engineering University of Surrey, UK  \nGeorgios Leontidis  \nDepartment of Computing Science & Interdisciplinary Institute University of Aberdeen, UK  \n[a.li.21@abdn. ac.uk](a.li.21@abdn. ac.uk)  \n[aiden. durrant@abdn. ac.uk](aiden. durrant@abdn. ac.uk)  \n[milan.markovic@abdn. ac.uk](milan.markovic@abdn. ac.uk)  \n[t.huang2@exeter. ac.uk](t.huang2@exeter. ac.uk)  \n[souvikk.kundu@intel. com](souvikk.kundu@intel. com)  \n[tianlong@cs. unc. edu](tianlong@cs. unc. edu)  \n[l.yin@surrey. ac.uk](l.yin@surrey. ac.uk)  \n[georgios.leontidis@abdn. ac.uk](georgios.leontidis@abdn. ac.uk)  \nReviewed on OpenReview: [https: // openreview. net/ forum? id= XX9JdOJD8R](https: // openreview. net/ forum? id= XX9JdOJD8R)  \nAbstract  \nPruning of deep neural networks has been an effective technique for reducing model size while preserving most of the performance of dense networks, crucial for deploying models on memory and power-constrained devices. While recent sparse learning methods have shown promising performance up to moderate sparsity levels such as 95% and 98%, accuracy quickly deteriorates when pushing sparsities to extreme levels due to unique challenges such as fragile gradient flow. In this work, we explore network performance beyond the commonly studied sparsities, and develop techniques that encourage stable training without accuracy collapse even at extreme sparsities, including 99 .90%, 99 .95% and 99 .99% on ResNet architectures. We propose three complementary techniques that enhance sparse training through different mechanisms: 1) Dynamic ReLU phasing, where DyReLU initially allows for richer parameter exploration before being gradually replaced by standard ReLU, 2) weight sharing which reuses parameters within a residual layer while maintaining the same number of learnable parameters, and 3) cyclic sparsity, where both sparsity levels and sparsity patterns evolve dynamically throughout training to better encourage parameter exploration. We eval-  \nuate our method, which we term Extreme Adaptive Sparse Training (EAST) at extremesparsities using ResNet-34 and ResNet-50 on CIFAR-10, CIFAR-100, and ImageNet, achieving competitive or improved performance compared to existing methods, with notable gains at extreme sparsity levels. Code is available at [https://github.com/TensorStrike/ES2](https://github.com/TensorStrike/ES2) .  \n1 Introduction  \nNetwork pruning (Han et al., 2015a;b; LeCun et al., 1990; Liu et al., 2017; Li et al., 2016; Kusupati et al. , 2020) is a widely-used technique for reducing a network’s parameters and compressing its size. Reducing model sizes is crucial for deploying models on edge devices with limited resources. Conventionally, pruning methods have focused on reducing parameters from pre-trained models. However, it requires at least as much computation as training a dense model as it must converge before pruning takes place. The Lottery Ticket Hypothesis work (Frankle & Carbin, 2018; Frankle et al., 2020a; Malach et al., 2020) gives theoretical foundation that subnetworks have the potential to reach full performance even when trained from an initially sparse state. This insight has recently gained much traction in sparse training, a paradigm in which sparse networks are trained from scratch without the need for dense pre-training.  \nSparse training methods can be broadly classified into two categories: static sparse training (SST), also somet","cbCainmFea0c6N54","https://ap.wps.com/l/cbCainmFea0c6N54","pdf",777183,1,15,"English","en",105,"# Introduction\n## Background on network pruning and sparse training\n## Static sparse training vs dynamic sparse training\n## Limits of sparse training and problem motivation\n## Proposed framework: Extreme Adaptive Sparse Training (EAST)","[{\"question\":\"What problem does this paper address about sparse training?\",\"answer\":\"Sparse training performs well up to moderate sparsity, but accuracy collapses at extreme sparsity due to challenges like fragile gradient flow and layer collapse. The paper investigates these limits and targets extreme sparsity regimes.\"},{\"question\":\"What is EAST and what techniques does it use?\",\"answer\":\"The paper proposes Extreme Adaptive Sparse Training (EAST), combining Dynamic ReLU phasing, weight sharing within residual layers, and cyclic sparsity scheduling to encourage stable learning at extreme sparsities.\"},{\"question\":\"On which models and datasets is the method evaluated?\",\"answer\":\"EAST is evaluated on ResNet-34 and ResNet-50 across CIFAR-10, CIFAR-100, and ImageNet, reporting competitive or improved performance especially at extreme sparsity levels.\"}]","Pushing the Limits of Sparsity - A Bag of Tricks for Extreme Pruning | PDF",1785811007,38,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"pushing-the-limits-of-sparsity-a-bag-of-tricks-for-extreme-pruning","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/pushing-the-limits-of-sparsity-a-bag-of-tricks-for-extreme-pruning/122509/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does this paper address about sparse training?","Question",{"text":75,"@type":76},"Sparse training performs well up to moderate sparsity, but accuracy collapses at extreme sparsity due to challenges like fragile gradient flow and layer collapse. The paper investigates these limits and targets extreme sparsity regimes.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What is EAST and what techniques does it use?",{"text":80,"@type":76},"The paper proposes Extreme Adaptive Sparse Training (EAST), combining Dynamic ReLU phasing, weight sharing within residual layers, and cyclic sparsity scheduling to encourage stable learning at extreme sparsities.",{"name":82,"@type":73,"acceptedAnswer":83},"On which models and datasets is the method evaluated?",{"text":84,"@type":76},"EAST is evaluated on ResNet-34 and ResNet-50 across CIFAR-10, CIFAR-100, and ImageNet, reporting competitive or improved performance especially at extreme sparsity levels.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]