Beyond Negative-Ridge Endpoints: Mixed-Sign Spectral Regularization via Negative-Shifted Gradient Descent

2026-07-24Machine Learning

Machine Learning
AI summary

The authors study a type of linear regression where there are more features than data points, focusing on how certain directions in the data act like a penalty term called ridge regression. They find that the usual negative ridge penalty has limitations, but an early-stopped gradient descent with a negative shift avoids these issues and achieves better control over the solution. Using a specific Gaussian data model, the authors show that this approach can outperform standard methods by balancing between shrinking large and small data directions differently. Their work involves advanced mathematical tools to handle the complex behavior of this shifted gradient process and ensures the method works well when selecting models based on validation data.

Overparameterized linear regressionRidge regressionNegative ridge penaltyGradient descentEarly stoppingEmpirical eigenvaluesMarchenko-Pastur lawSpike-plus-flat modelImplicit regularizationDuhamel integrals
Authors
Peng Zhao
Abstract
In overparameterized linear regression, many weak spectral directions act like a ridge penalty on the signal-bearing spectrum; negative ridge is the natural correction, pushing filters above one. The stable negative-ridge endpoint, however, is structurally limited: its pole must stay below the smallest nonzero empirical eigenvalue, and it anti-shrinks smaller eigenvalues more than larger ones. Early-stopped negative-shifted gradient descent escapes this constraint. Its filter is smooth at the would-be pole and mixed-sign-capable: above-ridgeless directions form a leading prefix, with lower directions shrunk or exposure-controlled while stopping sets the crossover. In a Gaussian spike-plus-flat model we discover a Marchenko-Pastur barrier: the shift that cancels the implicit penalty lies a bulk width above the smallest empirical eigenvalue, and the stopped path improves on every admissible endpoint by a polynomial factor in risk under explicit conditions. Our main theorem permits a general high-effective-rank tail: its trace sets the implicit floor, its squared spectrum controls exposure, and the floor-critical path recovers all head scales at once, beyond positive shrinkage and, once scales separate, every uniform rescaling of ridgeless. Handling the noncontractive shifted dynamics is the central technical challenge; localized Duhamel integrals control them. A finite-grid hold-out inequality transfers the separations to the validation-selected algorithm.