Tag
This paper analyzes the backward-pass geometry of Z-loss in AI model training, providing a framework for understanding and optimizing gradients in dense output heads and sparse mixture-of-experts routers, with evaluations on GPT-2 and Pythia models.