AI summaryⓘ
Shani and colleagues studied how large language models (LLMs) represent categories and typical examples but found previous methods missed some details. They used a different method called sparse autoencoder (SAE) sets to measure similarities, hoping it would be clearer. While SAE sets worked well in toy models and made semantic sense in text, the authors found they did not match human category boundaries or judgments about typical examples better than older methods. Instead, SAE sets seemed to show the model's own internal similarities. When they tested how SAE sets changed with clear concept shifts, these changes did not line up with how humans perceive conceptual changes, suggesting SAE features don't simply combine like basic parts outside of ideal cases.
large language modelscategory boundariestypicalitycosine similaritysparse autoencoderlatent representationsemantic compositionalitymodel interpretabilityconceptual change
Authors
Nikolai Bolik, Lennart Stöpler, Artur Andrzejak
Abstract
Shani et al. (2026) show that LLM representations broadly recover human category boundaries, while failing to reflect fine-grained typicality structure. Their analysis uses cosine similarity over dense model representations. We revisit their approach using overlap over active sparse autoencoder (SAE) latent sets as a more interpretable similarity measure. We first verify that this set-level measure is meaningful: SAE latent sets can recover union-like compositional structure in controlled toy models and induce semantically coherent neighborhoods in natural text. Extending the human-concepts analysis to SAE set similarities, we find that SAE activation sets do not recover human category boundaries or within-category typicality more faithfully than dense embeddings or residual-stream states, but instead track model-internal similarity structure. To probe this gap further, we study active latent sets under well-controlled semantic modifications, revealing a substantial mismatch between human judgements of conceptual change and change in the SAE active set. We interpret this as evidence that, outside idealised settings, SAE features do not compose via simple bag-of-features semantics.