@percyliang: Marin believes strongly in the power of the open community. There are so many datasets that have been produced, living …

X AI KOLs Timeline Models

Summary

Marin has built scalable infrastructure to process 25T tokens from open datasets on Hugging Face for training their 535B language model, emphasizing the power of the open community.

Marin believes strongly in the power of the open community. There are so many datasets that have been produced, living on Hugging Face. We have built scalable infrastructure that ingests this data, performs deduplication, decontamination, mixing, producing 25T tokens.
Original Article
View Cached Full Text

Cached at: 09/26/26, 03:05 PM

Marin believes strongly in the power of the open community. There are so many datasets that have been produced, living on Hugging Face. We have built scalable infrastructure that ingests this data, performs deduplication, decontamination, mixing, producing 25T tokens.

Will Held (@WilliamBarrHeld): There’s an enormous amount of open training data on Hugging Face. What does it take to make it work together 🤗?

For Marin’s 535B run, we built on 25T tokens from 152 datasets with licenses permitting training.

Here’s the work between downloading those and training a model 🧵

Similar Articles

marin-community/marin

GitHub Trending (daily)

Marin is an open-source research program and software platform dedicated to the transparent development of foundation models, covering data curation to model training and evaluation.