Tag
Researchers from Harvard and UIUC discovered a third pretraining axis that improves sample efficiency by 6.2x and speeds up GenAI generation by 250x.
Achieved 1000 tokens per second generation on Qwen3.6 27B using V100 GPUs with 128 concurrent requests, and 80 t/s for single user.