Tag
This paper presents an open-source pretraining recipe that trains 2B-parameter language models on consumer-grade RTX 5090 GPUs for under $7K, achieving performance near larger baselines and deriving cost scaling laws for model training.