Learning between the peaks: sharp asymptotics for kernel ridge regression under power-law anisotropy

2026-08-28Machine Learning

Machine Learning
AI summary

The authors study how kernel ridge regression learns when the input data has directions with different levels of variability (anisotropy), focusing on high-dimensional data with polynomial kernels. They find that when anisotropy is weak, learning behaves somewhat like the usual case but with some changes in how bias and variance evolve as more data is added. When anisotropy is strong, the problem effectively becomes low-dimensional, and variance stops improving with more data while bias shows sharp transitions depending on the target function's properties. They also explore how the alignment of the target function with the main data directions affects learning. Overall, the authors clarify how the shape of the input data influences kernel methods' performance.

kernel ridge regressionanisotropyGaussian datahigh-dimensional statisticspolynomial kernelbias-variance tradeofflearning curveseffective dimensionridgeless interpolationsingle-index model
Authors
Lorenzo Rizzi, Arie Wortsman Zurich, Bruno Loureiro
Abstract
We study kernel ridge regression under anisotropic Gaussian data, where the input covariance decays as a power law with exponent $α\geq 0$ for polynomial inner-product kernels. We derive asymptotically sharp expressions for the kernel spectrum and the generalization error in the polynomial high-dimensional regime $n=Θ(d^κ)$, revealing how anisotropy reshapes the learning curves. For weak anisotropy ($0<α<1$), the problem remains effectively high-dimensional and retains some features of the isotropic case, while departing from it in others: the variance still peaks at integer sample complexities $κ\in\mathbb{N}$, but these peaks are progressively damped as $α$ grows; meanwhile, for targets strongly aligned with the data's principal directions, the bias drops at fractional sample complexities, decoupling the bias transitions from the interpolation peaks. For strong anisotropy ($α> 1$), the effective dimension of the problem is constant, and the variance stops depending on sample size altogether, plateauing under ridgeless interpolation or vanishing at an explicit rate under fixed ridge penalty. The bias undergoes a sharp transition governed by the target's decay rate: below a threshold, learning is abrupt rather than gradual; above it, the bias decays as a power law that recovers the classical source and capacity rates. We finally specialize these results to single-index targets, showing how the alignment of the index with the data's principal directions determines the effect of anisotropy on learning. Together, our results clarify how the input geometry shapes the kernel features and fundamentally impacts its generalization properties.