Emergent Generalization by Representation Learning in Artificial Neural Networks

2026-07-11Information Theory

Information TheoryMachine LearningNeural and Evolutionary Computing
AI summary

The authors studied how simplifying complex brain activity into low-dimensional patterns helps in learning and predicting tasks. They found that forcing a neural network to create these simple patterns improves its ability to generalize to new situations and to remember over time. They also measured how these patterns change as learning progresses, showing a rise and fall pattern that relates to better performance. Similar brain activity patterns were observed in mice learning a maze task, suggesting these low-dimensional representations are important for flexible thinking. Overall, the authors suggest that such compact representations help both artificial and biological systems learn better.

Dimensionality ReductionNeural ManifoldsRecurrent Neural NetworksInformation BottleneckCausal EmergenceGeneralisationTime-Series PredictionHippocampusNeural CodingMemory and Learning
Authors
Hardik Rajpal, Dan Goodman
Abstract
Dimensionality reduction has proven powerful for identifying neural manifolds, which are low-dimensional structures underlying high-dimensional neural activity. These low-dimensional representations have improved the interpretability of population-level coding. Yet whether such low-dimensional representations are biologically relevant and confer functional advantages in learning systems, or merely reflect neuron-level activity, remains contested in neuroscience. We show that an explicit information bottleneck forcing a recurrent neural network to learn a low-dimensional representation is necessary for rotational and out-of-distribution generalisation in a time-series prediction task. Using information-theoretic measures of causal emergence, we characterise the dynamics of this representation across the memorisation-to-generalisation transition, finding a non-monotonic trajectory which shows an initial decrease, a minimum, and a subsequent rise to a maximum, even as prediction loss falls monotonically. This trajectory scales with task complexity, and the magnitude of emergent structure reliably predicts generalisation performance. Analysis of CA1 hippocampal activity in mice learning an alternating maze task reveals analogous non-monotonic emergence dynamics that track behavioural performance. Together, these findings indicate that the ability of neural networks to learn compact, distributed and emergent representations confers a functional advantage for generalisation, supporting a causal role for learned representations in cognition.