dots-studio/dots3-note-prev · Hugging Face
Summary
dots-studio releases dots3-note preview, the first open-weight model in the dots3 family: a 280B-parameter multimodal MoE with 16B activated parameters, 512K context, and support for text, image, video, and audio understanding.
View Cached Full Text
Cached at: 08/13/26, 11:30 PM
dots-studio/dots3-note-prev · Hugging Face
Source: https://huggingface.co/dots-studio/dots3-note-prev 中文| English

dots3-note Preview
🌐Tech Blog| 📄Full Report (coming soon)
https://huggingface.co/dots-studio/dots3-note-prev#table-of-contentsTable of Contents
- Model Introduction
- Model Overview
- Evaluation Results- General Reasoning and Agent - Multimodal Understanding
- Model Links
- Quickstart
- Deployment- Transformers - SGLang - vLLM
- Benchmark Appendix
- License
- Contact Us
https://huggingface.co/dots-studio/dots3-note-prev#model-introductionModel Introduction
dots3-note preview is the first open-weight model in the dots3 family. It is a Mixture-of-Experts model with 280B total parameters, 16B activated parameters, and support for a context length of up to 512K tokens. The model can understand text, images, video, and audio, and produces text outputs.
dots3-note preview is optimized for a broad range of tasks, including:
- General knowledge and instruction following;
- Mathematical and logical reasoning;
- Tool use and multi-step agent workflows;
- Interactive tasks that require exploration, memory updates, and adaptation;
- Code generation and code-based problem solving;
- Image, document, chart, audio, and video understanding;
- Long-context information processing.
The dots3 family is designed to include models with different trade-offs among capability, latency, and inference cost. dots3-note preview is the most lightweight member of the family.
https://huggingface.co/dots-studio/dots3-note-prev#model-overviewModel Overview
PropertyValueArchitectureMultimodal MoETotal Parameters280BActivated Parameters16BMTP1 shared layer, 1.13BNumber of Layers1 dense + 45 MoEHidden Size5120FFN Hidden Size13824 (dense), 1536 (per expert)Experts256 routed + 1 shared, top-8Attention13 DSA + 33 SWA (~1:3)DSATop-2048Context Length512KVocabulary Size152KVision EncoderMoE ViT, 7B total, 1.2B activatedAudio EncoderDense, 800MSupported PrecisionBF16, FP8InputText, image, video, audioOutputText
https://huggingface.co/dots-studio/dots3-note-prev#evaluation-resultsEvaluation Results
https://huggingface.co/dots-studio/dots3-note-prev#general-reasoning-and-agentGeneral Reasoning and Agent
https://huggingface.co/dots-studio/dots3-note-prev#multimodal-understandingMultimodal Understanding
https://huggingface.co/dots-studio/dots3-note-prev#model-linksModel Links
https://huggingface.co/dots-studio/dots3-note-prev#quickstartQuickstart
Recommended: serve the FP8 checkpoint on one 8-GPU node withSGLangorvLLM.
from openai import OpenAI
client = OpenAI(base_url="http://127.0.0.1:8000/v1", api_key="EMPTY")
response = client.chat.completions.create(
model="dots3-note-prev",
messages=[
{"role": "user", "content": "Hello! Can you briefly introduce yourself?"},
],
temperature=1.0,
top_p=0.95,
max_tokens=256,
# Set enable_thinking=True for reasoning; False returns a direct response.
extra_body={"chat_template_kwargs": {"enable_thinking": False}},
)
print(response.choices[0].message.content)
For a multimodal request, replacemessageswith one of these public examples:
examples = {
"image": [
{"type": "image_url", "image_url": {"url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/cats.png"}},
{"type": "text", "text": "How many cats are in this image?"},
],
"audio": [
{"type": "audio_url", "audio_url": {"url": "https://huggingface.co/datasets/hf-internal-testing/dummy-audio-samples/resolve/main/mary_had_lamb.mp3"}},
{"type": "text", "text": "Transcribe this nursery rhyme."},
],
"video": [
{"type": "video_url", "video_url": {"url": "https://huggingface.co/datasets/merve/vlm_test_images/resolve/main/concert.mp4"}},
{"type": "text", "text": "Describe the performance and what can be heard."},
],
}
messages = [{"role": "user", "content": examples["image"]}]
Video inputs include their audio track when available.
https://huggingface.co/dots-studio/dots3-note-prev#deploymentDeployment
The commands below target FP8 on one 8-GPU node. BF16 requires more memory. Tune the context length to available memory, concurrency, and input modalities.
Native support is available onvLLMmain.Transformers #47844andSGLang #33829are still under review; until they are merged, use the PR revisions below.
https://huggingface.co/dots-studio/dots3-note-prev#transformersTransformers
First install mutually compatiblePyTorch and torchvisionbuilds supported by your NVIDIA driver. For audio and video, also install a PyTorch-compatibletorchcodec(included below) and FFmpeg with your system package manager. Then installTransformers #47844:
pip install accelerate pillow torchcodec kernels==0.16.0 "transformers @ git+https://github.com/huggingface/transformers.git@refs/pull/47844/head"
Run a minimal local inference:
from transformers import AutoModelForMultimodalLM, AutoProcessor
model_id = "dots-studio/dots3-note-prev-fp8"
processor = AutoProcessor.from_pretrained(model_id)
model = AutoModelForMultimodalLM.from_pretrained(model_id, dtype="auto", device_map="auto")
messages = [
{"role": "user", "content": "Hello! Please briefly introduce yourself."},
]
inputs = processor.tokenizer.apply_chat_template(
messages,
add_generation_prompt=True,
return_tensors="pt",
return_dict=True,
enable_thinking=False,
).to(model.device)
outputs = model.generate(**inputs, max_new_tokens=128)
print(processor.decode(outputs[0, inputs.input_ids.shape[1] :], skip_special_tokens=True))
Use SGLang or vLLM for multi-GPU OpenAI-compatible serving.
https://huggingface.co/dots-studio/dots3-note-prev#sglangSGLang
Recommended: use the release imagelmsysorg/sglang:dev-dots3-note. Full one-node recipes and tuning notes are in theDots3-Note cookbook. Source support is tracked inSGLang #33829.
Docker (the image downloads the checkpoint from Hugging Face on first run):
docker run --gpus all --ipc=host -p 8000:8000 \
lmsysorg/sglang:dev-dots3-note \
sglang serve \
--model-path dots-studio/dots3-note-prev-fp8 \
--served-model-name dots3-note-prev \
--host 0.0.0.0 \
--port 8000 \
--context-length 524288 \
--enable-dp-attention \
--dp-size 8 \
--tp-size 8 \
--ep-size 8 \
--moe-dense-tp-size 1 \
--page-size 64 \
--trust-remote-code \
--attention-backend fa3 \
--moe-a2a-backend deepep \
--enable-multimodal \
--speculative-algorithm NEXTN \
--speculative-num-steps 3 \
--speculative-eagle-topk 1 \
--speculative-num-draft-tokens 4 \
--speculative-draft-model-path dots-studio/dots3-note-prev-fp8
Or install from source / the PR and run the samesglang servearguments locally.\-\-attention\-backend fa3sets prefill, decode, and (when speculative decoding is enabled) draft attention. MTP/NEXTN (\-\-speculative\-algorithm NEXTNand the related flags) is optional and can reduce TPOT by more than 50%. Prefill CUDA graph is not supported yet.
Optional features:
# Load only the language model
--language-only
# Enable OpenAI-compatible tool calling
--tool-call-parser dots
https://huggingface.co/dots-studio/dots3-note-prev#vllmvLLM
Native dots3-note preview support is available onvLLMmain. Use a recent nightly build until it is included in a stable release.
The following example deploys the FP8 checkpoint on eight NVIDIA H100 GPUs with TP=8 and EP=8:
vllm serve dots-studio/dots3-note-prev-fp8 \
--served-model-name dots3-note-prev \
--host 0.0.0.0 \
--tensor-parallel-size 8 \
--enable-expert-parallel \
--moe-backend deep_gemm \
--max-model-len 262144
Optional features:
# Load only the language model
--language-model-only
# Enable three-token MTP speculative decoding
--speculative-config '{"method":"mtp","num_speculative_tokens":3}'
# Enable OpenAI-compatible automatic tool calling
--enable-auto-tool-choice --tool-call-parser dots
https://huggingface.co/dots-studio/dots3-note-prev#benchmark-appendixBenchmark Appendix
https://huggingface.co/dots-studio/dots3-note-prev#licenseLicense
Copyright (c) 2026 Xiaohongshu.
Developed and released by dots studio.
The dots3-note preview model weights and modeling code in this repository are licensed under the Apache License, Version 2.0.
See the LICENSE file for details.
Transformers, SGLang, vLLM, and other third-party software are subject to their respective licenses.
https://huggingface.co/dots-studio/dots3-note-prev#contact-usContact Us
For questions and feedback, please contact us through:
- Email:[email protected]
dots3-note preview is developed and released by dots studio.
Similar Articles
dots3-note Preview targets autonomous agency and dynamic real-world environments with 512K context
dots3-note Preview is an AI model targeting autonomous agency and dynamic real-world environments with a 512K context window.
@AdinaYakup: More players are joining the open source summer RedNote just released dots3-note preview The first open weight model in…
RedNote has released dots3-note preview, the first open-weight model in the dots3 family, featuring 280B parameters with 16B active, multi-modal understanding (text, image, video, audio), 512K context length, and Apache 2.0 license, with strong agent capabilities.
dots.tts 2B🎙️ SOTA TTS from RedNote
RedNote releases dots.tts, a 2B parameter open-source text-to-speech model with zero-shot voice cloning and 48 kHz synthesis.
blokdots 3.0
Blokdots 3.0 is a product update for a hardware sketching tool, enabling users to sketch with hardware components.
@HuggingPapers: Google just released Magenta RealTime 2 on Hugging Face The only open-weights model for real-time continuous music gene…
Google released Magenta RealTime 2 on Hugging Face, an open-weights model for real-time continuous music generation on device with ~200ms latency, steerable by text, audio, or MIDI.



