Tag
An HFT veteran turned AI expert argues that AI winners will be companies that optimize token spend per dollar, not the biggest spenders, and compares token allocation to capital allocation on a trading floor.
Proposes a three-term scaling law that decouples model size, training steps, and batch size, enabling robust fitting with fewer runs and deriving scaling laws for suboptimal batch sizes.