Skip to content
← Publications

ICLR 2023 · Feb 2023

Understanding Edge-of-Stability Training Dynamics with a Minimalist Example

Xingyu Zhu*, Zixuan Wang*, Xiang Wang, Mo Zhou, Rong Ge

* Equal contribution

GD trajectory on EoS for a minimalist model
Figure. GD trajectory on EoS for a minimalist model

Abstract

Recently, researchers observed that gradient descent for deep neural networks operates in an “edge-of-stability” (EoS) regime: the sharpness (maximum eigenvalue of the Hessian) is often larger than stability threshold 2/η2/\eta (where η\eta is the step size). Despite this, the loss oscillates and converges in the long run, and the sharpness at the end is just slightly below 2/η2/\eta. While many other well-understood nonconvex objectives such as matrix factorization or two-layer networks can also converge despite large sharpness, there is often a larger gap between sharpness of the endpoint and 2/η2/\eta. In this paper, we study EoS phenomenon by constructing a simple function that has the same behavior. We give rigorous analysis for its training dynamics in a large local region and explain why the final converging point has sharpness close to 2/η2/\eta. Globally we observe that the training dynamics for our example has an interesting bifurcating behavior, which was also observed in the training of neural nets.

Citation

@misc{zhu_understanding_2023,
  author    = {Zhu, Xingyu and Wang, Zixuan and Wang, Xiang and Zhou, Mo and Ge, Rong},
  month     = {February},
  publisher = {arXiv},
  title     = {Understanding Edge-of-Stability Training Dynamics with a Minimalist Example},
  url       = {http://arxiv.org/abs/2210.03294},
  year      = {2023}
}