training-steps

Tag

Cards List
#training-steps

How to Allocate Your Tokens? Scaling Laws with Training Steps and Batch Size

arXiv cs.LG · 2026-07-03 Cached

Proposes a three-term scaling law that decouples model size, training steps, and batch size, enabling robust fitting with fewer runs and deriving scaling laws for suboptimal batch sizes.

0 favorites 0 likes
← Back to home

Submit Feedback