Weight-Adjusted Gradients Reveal Parameter Importance and Failure Modes in LLMs

2026-07-12Machine Learning

Machine LearningArtificial Intelligence
AI summary

The authors introduce a method called Weight-Adjusted Gradients (WAG) to figure out which parts of large language models are most important by looking at how model weights interact with gradients. They find that a small number of parameters have a big impact on the model’s behavior, especially causing failures that previous methods missed. Their work shows that combining weight and gradient information gives a better understanding of parameter importance than using either alone. They also demonstrate that WAG can be used in various practical tasks like adjusting experts in models, unlearning specific data, and editing knowledge within the model.

Large Language ModelsParameter ImportanceGradientsWeightsCollapse PhenomenaMixture-of-ExpertsUnlearningQuantizationKnowledge EditingModel Interpretability
Authors
Shrestha Datta, Hongfu Liu, Anshuman Chhabra
Abstract
Understanding which parameters are influential in Large Language Models (LLMs) is central to improving their efficiency, reliability, and interpretability. We introduce Weight-Adjusted Gradients (WAG), a simple yet effective approach for estimating parameter importance that explicitly captures the interaction between model weights and first-order gradient information and identifies parameters that disproportionately influence model behavior, such as those responsible for collapse phenomena in LLMs. Across a range of models and settings, we show that WAG surfaces a tiny but critical subset of parameters whose modification leads to dramatic degradation in performance, a failure mode that existing importance metrics overlook. These findings reveal a previously underexplored interplay between weights and gradients, suggesting that parameter importance cannot be fully understood through either signal alone. The surprising effectiveness of WAG points to fundamental structural properties of trained networks and motivates new open questions about the role of zeroth-order and first-order information in deep learning. We demonstrate the practical utility of WAG across multiple applications, including expert allocation in mixture-of-expert architectures, parameter-specific unlearning, mixed-precision quantization, and layer selection for knowledge editing. Our results position WAG as a unified approach for analyzing, debugging, and controlling LLMs, and opens new directions for principled model-level interpretation.