@ClementDelangue: The future of biology shouldn’t stay behind black-box APIs. Especially when it touches personal health. Whether you’re …

X AI KOLs Following Models

Summary

Hugging Face releases Carbon, an open-source DNA base model that is 275x faster than comparable models, enabling local processing of whole genomes on a single GPU.

The future of biology shouldn’t stay behind black-box APIs. Especially when it touches personal health. Whether you’re @bryan_johnson measuring every biomarker, or @sytses openly sharing and analyzing his own immune-genetics data, you need open, local, transparent AI. @huggingface wasn’t created to be a biology company. It’s not the most obvious focus for us. But it feels too important not to do something. That’s why we built and released Carbon : a frontier DNA base model with open weights, training code and data pipeline, designed to be fine-tuned or continually pretrained for downstream biological tasks. Carbon is 275x faster than the next best model at its size. Fast enough to run locally on your laptop. Powerful enough to process a whole human genome on a single GPU in less than 2 days. The technical unlock: a DNA-native tokenizer that splits sequences into 6-base chunks for efficiency, while preserving single-base resolution during training and inference. More people able to inspect, run, fine-tune, improve and build on top of the models shaping biology. Open weights: https://huggingface.co/collections/HuggingFaceBio/carbon… Dataset: https://huggingface.co/datasets/HuggingFaceBio/carbon-pretraining-corpus… Demo: https://huggingface.co/spaces/HuggingFaceBio/carbon-demo… Let's go open AI biology!
Original Article
View Cached Full Text

Cached at: 05/20/26, 12:30 PM

The future of biology shouldn’t stay behind black-box APIs. Especially when it touches personal health.

Whether you’re @bryan_johnson measuring every biomarker, or @sytses openly sharing and analyzing his own immune-genetics data, you need open, local, transparent AI.

@huggingface wasn’t created to be a biology company. It’s not the most obvious focus for us. But it feels too important not to do something.

That’s why we built and released Carbon : a frontier DNA base model with open weights, training code and data pipeline, designed to be fine-tuned or continually pretrained for downstream biological tasks.

Carbon is 275x faster than the next best model at its size. Fast enough to run locally on your laptop. Powerful enough to process a whole human genome on a single GPU in less than 2 days.

The technical unlock: a DNA-native tokenizer that splits sequences into 6-base chunks for efficiency, while preserving single-base resolution during training and inference. More people able to inspect, run, fine-tune, improve and build on top of the models shaping biology.

Open weights: https://huggingface.co/collections/HuggingFaceBio/carbon… Dataset: https://huggingface.co/datasets/HuggingFaceBio/carbon-pretraining-corpus… Demo: https://huggingface.co/spaces/HuggingFaceBio/carbon-demo…

Let’s go open AI biology!


🧬 Carbon - a HuggingFaceBio Collection

Source: https://huggingface.co/collections/HuggingFaceBio/carbon

HuggingFaceBio’s Collections

Similar Articles

Carbon: Decoding the Language of Life

Reddit r/LocalLLaMA

Hugging Face released Carbon, a family of open DNA foundation models that matches state-of-the-art performance of Evo2-7B while being 275x faster, using 6-mer tokenization, factorized loss, and curated genomic data.

@charles_irl: bits to bio

X AI KOLs Timeline

Modal announces integration with Claude Science, providing elastic compute for life sciences researchers with up to $100K in committed resources.