Tag
The paper analyzes the delayed effects of minibatch perturbations in AdamW by modeling it as a finite-horizon input-state-output system, revealing how optimizer states influence training dynamics.