[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86405-en":3,"doc-seo-86405-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86405,3848291630094,"Emma Wilson","https://eur-avatar.wpscdn.com/davatar_085a072bc5b1113ac321206ff7593b45",8,"Research & Report","Prune, Update and Trim Robust Structured Pruning for Large Language Models","Large Language Models (LLMs) are costly to run during inference, especially with long-context inputs or on resource-limited devices. Post-training pruning (PTP) reduces the model’s parameter and compute requirements by removing weights with limited impact on performance. Existing PTP methods prune FFN nodes and entire attention layers using fixed importance criteria. Putri improves this by updating unpruned FFN weights to offset pruning error, pruning FFN layers sequentially with prior updates, and pruning individual attention heads instead of whole layers, extending to Grouped-Query Attention.","Prune, Update and Trim: Robust Structured Pruning for Large Language Models  \nDiego Coello de Portugal Mecke∗  \nISMLL & DARC VWFS University of Hildesheim Hildesheim, Germany  \nTom Hanika  \nISMLL University of Hildesheim  \nHildesheim, Germany  \nLars Schmidt-Thieme  \nISMLL University of Hildesheim  \nHildesheim, Germany  \narXiv :2605 . 1833 1v2 [ cs .LG] 13 Jul 2026  \nAbstract  \nLarge Language Models (LLMs) have experienced significant growth and development in recent years. However, performing inference on LLMs remains costly, especially for long-context inference or in resource-constrained devices. This motivates the development of new post-training pruning (PTP) methods. These methods reduce LLMs’ requirements by removing a substantial part of the model’s parameters. The discarded weights are selected depending on their impact on the models performance. Current PTP methods prune the models by removing the less informative hidden nodes from the FFN layers, and the least important attention layers. We propose Putri, a PTP method that introduces three changes to the Stateof-the-art. First, we update the un-pruned weights of the FFN to compensate for the introduced pruning error. Second, the FFN layers are pruned sequentially, taking into account the updates done to the previous layers. Third, instead of removing full attention layers, we remove individual attention-heads. We extend this method such that it can also address Grouped-Query Attention. In summary, Putri is a structure pruning method which remains simple while showing SOTA performance.  \nPruning experiments on multiple models with a wide variety of sparsity ranges and on different datasets, validate the generality of Putri. Notably, we demonstrate that, unlike previous methods, Putri can prune LLMs on extreme sparsity ratios. The code is available at: [https://github.com/Coello-dev/Putri](https://github.com/Coello-dev/Putri).  \n1 Introduction  \nLarge Language Models (LLMs) [1, 2, 3] show consistent improvements when increasing model scale, training data and compute. This enhanced performance has allowed LLMs to tackle a variety of tasks such as agentic workflow [4, 5, 6] and robotics [7, 8, 9] . Nevertheless, these tasks require the model to either run on a large number of tokens, which drastically increases the required memory, or run on a smaller device with less memory and computing capacity than data centers. These requirements have motivated the development of multiple post-training techniques that reduce the memory and computing requirements of LLMs by compressing model size.  \nFor example, Knowledge distillation [10] aims to train a smaller model (student) that mimics the prediction of the already trained bigger model (teacher) . Even though Knowledge distillation achievesa highly optimized smaller model, it requires the computational effort to run inference on the teacher model and train the student model at the same time. Another approach is Quantization [11], which reduces the memory requirements by reducing the floating-point precision of the model’s parameters. However, this technique is limited by the number of precision bits to store the weights. This sets a lower bound of how much it can compress the model. Another prominent method is Pruning [12], which aims to remove the less informative parameters of a model while maintaining the original model  \n∗ coellod@uni-hildesheim .de  \nPreprint.  \nperformance. This allows Pruning to reduce the model size with a comparatively small computational effort (with respect to Knowledge Distillation) . At the same time it is (in theory) capable of reducing the model size as much as required for a certain task (in contrast to Quantization) . Nevertheless, in practice Pruning methods are not used due to its inability to meaningfully reduce the model’s size.  \nThe first Pruning methods aimed to remove the model’s weights without considering the model’s structure (Unstructured pruning [13, 14]) . These methods result in m","cbCail0Ug10d3vHp","https://ap.wps.com/l/cbCail0Ug10d3vHp","pdf",1530377,4,1,16,"English","en",105,"# Introduction\n## Motivation for Post-Training Pruning\n## Alternatives: Distillation and Quantization\n## Unstructured vs. Structured Pruning\n## Prior Work and Key Limitations\n# Proposed Method: Putri\n## Update Mechanism for FFN\n## Sequential FFN Pruning\n## Attention Head-Level Pruning\n## Extension to Grouped-Query Attention\n# Experiments and Results\n## Validation Across Models, Sparsity, and Datasets","[{\"question\":\"What problem does the document address in large language model deployment?\",\"answer\":\"Inference with LLMs is expensive in memory and compute, particularly for long-context settings or on devices with limited resources. This motivates post-training compression methods.\"},{\"question\":\"What are the main limitations of existing post-training pruning methods?\",\"answer\":\"Prior methods typically prune without updating remaining weights, which can cause reconstruction error at high sparsity. They also often use coarse granularity, such as removing whole attention modules.\"},{\"question\":\"How does Putri differ from state-of-the-art structured pruning approaches?\",\"answer\":\"Putri updates unpruned FFN weights to compensate pruning error, prunes FFN layers sequentially while accounting for earlier updates, and prunes individual attention heads rather than full attention layers. It also extends to Grouped-Query Attention and supports extreme sparsity ratios.\"}]",1784211547,40,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"prune-update-and-trim-robust-structured-pruning-for-large-language-models","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/prune-update-and-trim-robust-structured-pruning-for-large-language-models/86405/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-26","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the document address in large language model deployment?","Question",{"text":75,"@type":76},"Inference with LLMs is expensive in memory and compute, particularly for long-context settings or on devices with limited resources. This motivates post-training compression methods.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What are the main limitations of existing post-training pruning methods?",{"text":80,"@type":76},"Prior methods typically prune without updating remaining weights, which can cause reconstruction error at high sparsity. They also often use coarse granularity, such as removing whole attention modules.",{"name":82,"@type":73,"acceptedAnswer":83},"How does Putri differ from state-of-the-art structured pruning approaches?",{"text":84,"@type":76},"Putri updates unpruned FFN weights to compensate pruning error, prunes FFN layers sequentially while accounting for earlier updates, and prunes individual attention heads rather than full attention layers. It also extends to Grouped-Query Attention and supports extreme sparsity ratios.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,119,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":29,"slug":118},7,"Healthcare","healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]