Singular perturbations and hierarchical learning in two-layer neural networks
2026-07-12 • Machine Learning
Machine Learning
AI summaryⓘ
The authors study how a very wide two-layer neural network learns a simplified model in high dimensions when the model is not perfectly specified. They focus on how the two layers learn at different speeds and prove that certain simple parts of the model (constant and linear) are learned at exact predicted times. They also explore when the network starts to learn more complex quadratic parts and show that earlier learned parts affect later learning stages. Their proofs use advanced math about flows near special geometric shapes and show that some neurons gain much importance while others adjust as learning progresses.
population gradient flowtwo-layer neural networkmisspecified modelsingle-index modelhierarchical learningsingular perturbationintegral constraintsempirical measureneurons dynamicshigh-dimensional learning
Authors
Cédric Gerbelot, Jean-Christophe Mourrat
Abstract
We study the population gradient flow of an infinitely wide two-layer neural network learning a misspecified single-index model in high dimension. The two layers are optimized jointly, with a perturbative parameter tuning the relative training speed between the first and second layer. This setting was considered by Berthier, Montanari and Zhou in \cite{berthier2024learning}, who conjectured a hierarchical learning scenario with explicit timescales as the second layer is trained faster than the first. In this paper, we prove that the constant and linear components of the hidden link function are indeed recovered within the predicted timescales, at sharp explicit thresholds. We then analyze the onset of learning of the quadratic component and show that the components learned at earlier stages continue to influence the dynamics in an essential way. Our proof is based on quantitative approximation results for singularly perturbed flows evolving near a manifold defined by integral constraints. At a phenomenological level, we also show that the empirical measure of the weights displays singular behaviour when reaching the quadratic component of the hidden link, with a small fraction of neurons growing significantly while the remaining ones rearrange to preserve the components already learned.