@percyliang: Marin 535B-A23B started training this week! As usual, the whole process is open. Voyage plan: pretraining (80%) + midtr…
Summary
Training of the Marin 535B-A23B AI model has begun with an open process, involving pretraining and midtraining on 18.75T tokens using GB200 NVL72 hardware over about 3 months.
View Cached Full Text
Cached at: 08/22/26, 11:34 PM
Marin 535B-A23B started training this week! As usual, the whole process is open. Voyage plan: pretraining (80%) + midtraining (20%) on 18.75T tokens on 11 x GB200 NVL72 for ~3 months (2.7e24 FLOPs). Post-training will follow. Before kicking off the run, we trained a 4-rung scaling ladder from 1.6B-A61M (48B tokens) to 27.7B-A1.2B (926B tokens) to debug issues, and to make a forecast of our hero run. This is by far our biggest run, so definitely expecting the unexpected.
Similar Articles
@percyliang: For the next Marin model, we are putting together a new data mix. Currently we have 18T tokens, but could use more. So …
Percy Liang announces that for the next Marin model, they are compiling a new data mix and request high-quality token data for pre-training, mid-training, and SFT.
@timodonnell: Want to watch a 535B parameter (23B active) LLM get trained live? Follow along here https://wandb.ai/marin-community/ma…
A tweet announces the live training of a 535B parameter (23B active) large language model, with links to follow the process on Weights & Biases and GitHub.
@eliebakouch: one of my favorite projects is Marin from the stanford folks, they have a scientific approach to training, are ready to…
Marin is an open-source framework from Stanford for reproducible foundation model research, covering data curation, tokenization, training, and evaluation; it was used to train an 8B parameter model that outperforms Llama 3.1 8B.
@percyliang: Not only do we want to train a good model, we want to know it'll be good before we even start training. About a month a…
The Marin team pre-registered a predicted loss of 2.252 for a 129B parameter MoE model training run, and the actual result landed at 2.234, demonstrating accurate loss prediction before training.
@WilliamBarrHeld: To train better open models, we need predictable scaling. Delphi is Marin’s first step: we pretrained many small models…
Marin AI researchers, led by William Barr Held, introduce Delphi, a methodology that pretrains small models to accurately predict the training outcomes of larger 25B-parameter runs. This research aims to establish predictable scaling for more efficient open-source AI model development.