Interpretable AI with Local Distillation

2026-08-24Machine Learning

Machine Learning
AI summary

The authors present a method called local distillation that helps make complex AI models easier to understand without losing accuracy. They use a simple linear model near each prediction point, guided by a more complex 'teacher' model to decide which data points are most relevant. They add some randomness to check which features are reliably important and group similar patients based on these local models. Their approach works well on many datasets and reveals patient differences that traditional models miss.

local linear modelinggradient-boosted ensemblesblack-box modelsmodel interpretabilitylasso penaltyGaussian randomizationfeature selectionlocal distillationgene expression data
Authors
Erin Craig, Yiling Huang, Snigdha Panigrahi
Abstract
Modern AI models such as tabular foundation models and gradient-boosted ensembles can outpredict classical methods, but provide little basis for reasoning about their predictions. High-stakes decisions call for models that are both accurate and interpretable as built. Local linear modeling offers a path forward: a smooth regression function is locally well approximated by a linear one, allowing a linear fit near each query point to achieve high accuracy without sacrificing transparency. The challenges lie in learning what is "local" and developing statistical tools for interpretation. Here, we propose local distillation, in which a black-box "teacher" guides a regularized linear "student" model at each query point. The teacher (1) defines locality by upweighting training observations with similar predicted outcomes, and (2) anchors the fit with its prediction at the query point, included as a pseudo-observation whose weight is estimated from the data. For interpretation, we add a small amount of Gaussian randomization to the local objective and use refits to assess stability: selection frequencies identify reliable features at a query point, and clustering the randomized fits identifies stable subgroups across the data. Under the lasso penalty, we prove that this randomization yields feature-selection probabilities that are stable under small perturbations of the training responses. Across 17 benchmark datasets, local distillation nearly matches its AI teacher's accuracy while producing a sparse linear model at each test point. In a high-dimensional cancer gene expression example, the framework identifies patient subgroups whose local models use different genes; this heterogeneity is invisible to a global linear model, and difficult to surface in a black-box model.