On the Convergence of Adam, Revisited
2026-07-03 • Machine Learning
Machine Learning
AI summaryⓘ
The authors demonstrate that the projected Adam optimization algorithm, when using any values for its moment decay parameters, can sometimes perform poorly by maintaining a nonzero average regret, meaning it doesn't learn perfectly over time. This finding builds on previous work which only showed this problem under more restricted conditions. They use a specific sequence of simple mathematical functions to prove their point. Their results also apply to many popular Adam variants and a related random setting.
Projected AdamOnline optimizationMoment decay parametersAverage regretReddi-Kale-Kumar resultThree-periodic sequenceAdam variantsRegret bound
Authors
Steven Heilman, Sampad Mohanty
Abstract
We show that projected Adam for online optimization with arbitrary moment decay parameters $β_1,β_2\in[0,1)$ can have average regret bounded away from zero. A similar result of Reddi-Kale-Kumar from 2018 required $β_1<\sqrt{β_2}$. Similar to their result, we use a three-periodic sequence of linear functions on $[-1,1]$ with slopes $c,-1,-1$, though we use $c$ slightly larger than $2$. This nonzero average regret result extends to Adam variants such as AdamW, RMSProp, NAdam, Adan, AdaMax, Muon, and to an i.i.d. variant of the three-periodic sequence of slopes for Adam.