@percyliang: Marin believes strongly in the power of the open community. There are so many datasets that have been produced, living …
Summary
Marin has built scalable infrastructure to process 25T tokens from open datasets on Hugging Face for training their 535B language model, emphasizing the power of the open community.
View Cached Full Text
Cached at: 09/26/26, 03:05 PM
Marin believes strongly in the power of the open community. There are so many datasets that have been produced, living on Hugging Face. We have built scalable infrastructure that ingests this data, performs deduplication, decontamination, mixing, producing 25T tokens.
Will Held (@WilliamBarrHeld): There’s an enormous amount of open training data on Hugging Face. What does it take to make it work together 🤗?
For Marin’s 535B run, we built on 25T tokens from 152 datasets with licenses permitting training.
Here’s the work between downloading those and training a model 🧵
Similar Articles
@WilliamBarrHeld: There’s an enormous amount of open training data on Hugging Face. What does it take to make it work together ? For Mari…
The article discusses the process of utilizing open training data from Hugging Face to train Marin’s 535B model, which involved 25T tokens from 152 datasets with permissible licenses.
@AndrewYNg: In the fight to defend openness in AI, the Marin project is a precious demonstration of openness in model training, wit…
Andrew Ng highlights the Marin project as a valuable example of openness in AI model training, with Percy Liang announcing the start of training for Marin 535B-A23B using open code, data, and processes.
@percyliang: For the next Marin model, we are putting together a new data mix. Currently we have 18T tokens, but could use more. So …
Percy Liang announces that for the next Marin model, they are compiling a new data mix and request high-quality token data for pre-training, mid-training, and SFT.
@eliebakouch: one of my favorite projects is Marin from the stanford folks, they have a scientific approach to training, are ready to…
Marin is an open-source framework from Stanford for reproducible foundation model research, covering data curation, tokenization, training, and evaluation; it was used to train an 8B parameter model that outperforms Llama 3.1 8B.
marin-community/marin
Marin is an open-source research program and software platform dedicated to the transparent development of foundation models, covering data curation to model training and evaluation.