[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86041-en":3,"doc-seo-86041-105":30,"detail-sidebar-cat-0-en-105":96},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86041,1099514067438,"River Wang","https://ap-avatar.wpscdn.com/avatar/100002539ee87300030?x-image-process=image/resize,m_fixed,w_180,h_180&k=1780474512215547542",8,"Research & Report","Weight-Adjusted Gradients Reveal Parameter Importance and Failure Modes in LLMs","Understanding which parameters matter in Large Language Models (LLMs) is key for better efficiency, reliability, and interpretability. Weight-Adjusted Gradients (WAG) estimates parameter importance by combining model weight magnitude with first-order gradient information, highlighting parameters that disproportionately steer behavior. Across multiple model settings, WAG surfaces a very small but critical subset whose perturbation causes dramatic performance collapse that existing metrics miss. The results expose a structural weight–gradient interaction and enable debugging and control in expert allocation, unlearning, quantization, and knowledge editing.","arXiv :2607 . 10803v 1 [ cs .LG] 12 Jul 2026  \nWEIGHT-ADJUSTED GRADIENTS REVEAL PARAMETER IMPORTANCE AND FAILURE MODES IN LLMS  \nShrestha Datta 1 , Hongfu Liu2 , and Anshuman Chhabra 1*  \n1University of South Florida, Tampa, FL, USA  \n2Brandeis University, Waltham, MA, USA  \n* Corresponding Author  \n{shresthadatta, [anshumanc}@usf.edu](anshumanc}@usf.edu) , [hongfuliu@brandeis.edu](hongfuliu@brandeis.edu)  \nABSTRACT  \nUnderstanding which parameters are influential in Large Language Models (LLMs) is central to improving their efficiency, reliability, and interpretability. We introduce Weight-Adjusted Gradients (WAG), a simple yet effective approach for estimating parameter importance that explicitly captures the interaction between model weights and first-order gradient information and identifies parameters that disproportionately influence model behavior, such as those responsible for collapse phenomena in LLMs. Across a range of models and settings, we show that WAG surfaces a tiny but critical subset of parameters whose modification leads to dramatic degradation in performance, a failure mode that existing importance metrics overlook. These findings reveal a previously underexplored interplay between weights and gradients, suggesting that parameter importance cannot be fully understood through either signal alone. The surprising effectiveness of WAG points to fundamental structural properties of trained networks and motivates new open questions about the role of zeroth-order and first-order information in deep learning. We demonstrate the practical utility of WAG across multiple applications, including expert allocation in mixture-of-expert architectures, parameterspecific unlearning, mixed-precision quantization, and layer selection for knowledge editing. Our results position WAG as a unified approach for analyzing, debugging, and controlling LLMs, and opens new directions for principled model-level interpretation.  \n1 Introduction  \nLarge Language Models (LLMs) have demonstrated strong performance across a wide range of tasks, becoming a central component in modern machine learning systems [6, 53, 44] . While much of the progress in this area has focused on scaling model size, improving training procedures, and designing new architectures [30, 28, 41], comparatively less attention has been given to understanding the internal structure of trained models [15] . In particular, identifying which parameters are most influential to model behavior remains an important yet underexplored problem. Such understanding is critical for improving model reliability, enabling efficient adaptation, and supporting interpretability.  \nA natural way to study this problem is through parameter importance estimation. Intuitively, the importance of a parameter can be measured by evaluating how sensitive the model’s behavior is to perturbations of that parameter. Existing approaches typically rely on either parameter magnitudes or gradient-based signals. Magnitude-based methods, such as those used in pruning, assume that parameters with larger absolute values are more important [24, 18] . Gradient-based approaches instead measure sensitivity through first-order information, for example using saliency or influence approximations [32, 42, 54] . Second-order methods further incorporate curvature information, but often incur high computational cost and rely on approximations that may not scale well to modern LLMs [27, 17, 3] . Despite these efforts, the underlying relationship between weight and gradient information remains poorly understood, and existing methods offer little insight into how this interaction relates to severe failure modes such as extreme performance degradation. Specifically, existing methods often fail to consistently identify parameters that induce significant performance degradation when perturbed, particularly in large-scale models.  \nIn this paper, we examine parameter importance from the perspective of the interaction between p","cbCairoCRw2ip2oA","https://ap.wps.com/l/cbCairoCRw2ip2oA","pdf",3388802,5,1,21,"English","en",105,"# Introduction\n# Parameter Importance Estimation\n## Magnitude-based, Gradient-based, and Second-order Methods\n# Weight-Adjusted Gradients (WAG)\n## Efficient Computation and Core Insight\n# Results and Failure Modes\n# Practical Applications\n## Mixture-of-Experts and Quantization\n## Unlearning and Knowledge Editing","[{\"question\":\"What problem does Weight-Adjusted Gradients (WAG) address in LLMs?\",\"answer\":\"WAG targets the underexplored question of which specific parameters most influence LLM behavior, especially when perturbations lead to severe degradation and performance collapse that existing importance metrics overlook.\"},{\"question\":\"How does WAG estimate parameter importance?\",\"answer\":\"WAG combines zeroth-order information (parameter weights) with first-order gradient signals to quantify how strongly each parameter influences model behavior through their interaction, without relying on second-order curvature or retraining.\"},{\"question\":\"What failure mode does WAG help detect that other metrics miss?\",\"answer\":\"WAG identifies a tiny subset of parameters whose modification can trigger abrupt performance collapse, revealing a severe failure mode that magnitude-only or gradient-only importance measures fail to consistently surface.\"},{\"question\":\"Where can WAG be applied in practice according to the document?\",\"answer\":\"WAG is demonstrated for expert allocation in mixture-of-experts, mixed-precision quantization, targeted unlearning, and layer/submodule selection for knowledge editing, using the influential-parameter signal to guide control and debugging.\"}]",1784208032,53,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":91,"head_meta":93,"extra_data":95,"updated_unix":28},"weight-adjusted-gradients-reveal-parameter-importance-and-failure-modes-in-llms","",{"@graph":36,"@context":90},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/weight-adjusted-gradients-reveal-parameter-importance-and-failure-modes-in-llms/86041/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-26","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82,86],{"name":73,"@type":74,"acceptedAnswer":75},"What problem does Weight-Adjusted Gradients (WAG) address in LLMs?","Question",{"text":76,"@type":77},"WAG targets the underexplored question of which specific parameters most influence LLM behavior, especially when perturbations lead to severe degradation and performance collapse that existing importance metrics overlook.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does WAG estimate parameter importance?",{"text":81,"@type":77},"WAG combines zeroth-order information (parameter weights) with first-order gradient signals to quantify how strongly each parameter influences model behavior through their interaction, without relying on second-order curvature or retraining.",{"name":83,"@type":74,"acceptedAnswer":84},"What failure mode does WAG help detect that other metrics miss?",{"text":85,"@type":77},"WAG identifies a tiny subset of parameters whose modification can trigger abrupt performance collapse, revealing a severe failure mode that magnitude-only or gradient-only importance measures fail to consistently surface.",{"name":87,"@type":74,"acceptedAnswer":88},"Where can WAG be applied in practice according to the document?",{"text":89,"@type":77},"WAG is demonstrated for expert allocation in mixture-of-experts, mixed-precision quantization, targeted unlearning, and layer/submodule selection for knowledge editing, using the influential-parameter signal to guide control and debugging.","https://schema.org",{"og:url":52,"og:type":92,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":94,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":97},[98,102,106,110,114,119,124,127,132,135,139],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":111,"show_sort_weight":112,"slug":113},"Comic",60,"comic",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},6,"Technology",50,"technology",{"id":120,"doc_module":4,"doc_module_name":46,"category_name":121,"show_sort_weight":122,"slug":123},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":125,"slug":126},30,"research-report",{"id":128,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":130,"slug":131},9,"Religion & Spirituality",20,"religion-spirituality",{"id":130,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":130,"slug":134},"World Cup","world-cup",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":136,"slug":138},10,"Lifestyle","lifestyle",{"id":140,"doc_module":4,"doc_module_name":46,"category_name":141,"show_sort_weight":20,"slug":142},19,"General","general"]