Tag
A detailed write-up on training a 210M text-to-image diffusion transformer from scratch on a single GPU, sharing key measurements on attention sinks, loss as a health signal, and timestep shifting benefits.
OpenAI launches GPT-6 Astra, an AI model trained on over 100,000 GPUs at the Texas Stargate site.
A new feature called OpenResearch allows reproducing and experimenting on papers, with a one-click template to train Vector Policy Optimization (VPO) on ToolRL, enabling diverse answer generation and improved test-time search.
Andrej Karpathy open-sourced an autonomous research agent that runs its own ML experiments overnight using a single GPU, automatically iterating on improvements by editing code and keeping changes that lower validation loss.