Finding and using interpretable latents in a neutrino foundation model with sparse autoencoders
2026-08-26 • Artificial Intelligence
Artificial IntelligenceMachine Learning
AI summaryⓘ
The authors used a special technique called sparse autoencoder-based mechanistic interpretability to understand how a machine learning model trained on neutrino data works inside. They found a clear map of physical concepts the model learned, but noticed the part of the model predicting neutrino direction wasn’t using this useful information much. So, they trained another part to predict the model’s uncertainty in direction, which did rely on the learned physics and greatly improved accuracy. This shows their approach can uncover hidden physics knowledge in models and guide better ways to use it.
sparse autoencodermechanistic interpretabilityneutrinoIceCubefoundation modeldirection reconstructioncausal interventionangular resolutionuncertainty estimationlatent representation
Authors
Raphaël Bonnet-Guerrini, Johann Ioannou-Nikolaides, Inar Timiryasov, Vincenzo Piuri
Abstract
We present a first application of sparse-autoencoder-based mechanistic interpretability to particle physics. Studying a neutrino foundation model pretrained on IceCube data and fine-tuned for direction reconstruction, we identify a validated atlas of physical concepts in the model representation, using a strict validation protocol consisting of held-out tests, matched nuisance controls, and replication across independent dictionary trainings. Causal interventions show that the direction head barely draws on this atlas. Motivated by this underused information, we train an uncertainty head on the same event-level representation to predict the model's angular reconstruction error. Unlike the direction head, it depends causally on quality and brightness features from the atlas. At $20\%$ selection efficiency, this interpretable estimator improves the median angular resolution from $20.2^\circ$ to $3.2^\circ$. These results suggest that mechanistic interpretability can reveal learned latent physics encoded within a model's internal representation and help design downstream tasks that exploit it.