I trained a 0.5M model on 1B tokens of Fineweb-edu dataset.

Reddit r/LocalLLaMA Models

Summary

A personal project where a 0.5M parameter language model was trained on 1 billion tokens from the Fineweb-edu dataset.

No content available
Original Article

Similar Articles

[NEW] Supra-50M Released!

Reddit r/LocalLLaMA

SupraLabs released Supra-50M, a compact 50M-parameter causal language model with base and instruct versions, trained on 20B tokens from fineweb-edu, achieving competitive benchmarks against larger models like GPT-2 and SmolLM.