z-lab/Qwen3.6-35B-A3B-DFlash
Summary
z-lab releases DFlash, a speculative decoding drafter that uses a lightweight block-diffusion model to draft 15–16 tokens in parallel, yielding up to 2.9× speedup for Qwen3.6-35B-A3B inference.
View Cached Full Text
Cached at: 04/21/26, 07:47 PM
z-lab/Qwen3.6-35B-A3B-DFlash · Hugging Face
Source: https://huggingface.co/z-lab/Qwen3.6-35B-A3B-DFlash Paper|GitHub|Blog
DFlashis a speculative decoding method that uses a lightweightblock diffusionmodel to draft multiple tokens in parallel. This is the drafter model, which must be paired withQwen/Qwen3.6-35B-A3B.

https://huggingface.co/z-lab/Qwen3.6-35B-A3B-DFlash#quick-startQuick Start
https://huggingface.co/z-lab/Qwen3.6-35B-A3B-DFlash#installationInstallation
vLLM:
uv pip install vllm
uv pip install -U vllm --torch-backend=auto --extra-index-url https://wheels.vllm.ai/nightly
SGLang:
uv pip install "git+https://github.com/sgl-project/sglang.git@refs/pull/20547/head#subdirectory=python"
https://huggingface.co/z-lab/Qwen3.6-35B-A3B-DFlash#launch-serverLaunch Server
vLLM:
vllm serve Qwen/Qwen3.6-35B-A3B \
--speculative-config '{"method": "dflash", "model": "z-lab/Qwen3.6-35B-A3B-DFlash", "num_speculative_tokens": 15}' \
--attention-backend flash_attn \
--max-num-batched-tokens 32768
SGLang:
# Optional: enable schedule overlapping (experimental, may not be stable)
# export SGLANG_ENABLE_SPEC_V2=1
# export SGLANG_ENABLE_DFLASH_SPEC_V2=1
# export SGLANG_ENABLE_OVERLAP_PLAN_STREAM=1
python -m sglang.launch_server \
--model-path Qwen/Qwen3.6-35B-A3B \
--speculative-algorithm DFLASH \
--speculative-draft-model-path z-lab/Qwen3.6-35B-A3B-DFlash \
--speculative-num-draft-tokens 16 \
--tp-size 1 \
--attention-backend fa3 \
--mem-fraction-static 0.75 \
--mamba-scheduler-strategy extra_buffer \
--trust-remote-code
**Tip:**For long-context or agentic workloads, add
\-\-speculative\-dflash\-draft\-window\-size WINDOW\_SIZEto enable sliding-window attention for the drafter.
https://huggingface.co/z-lab/Qwen3.6-35B-A3B-DFlash#usageUsage
from openai import OpenAI
client = OpenAI(base_url="http://localhost:30000/v1", api_key="EMPTY")
response = client.chat.completions.create(
model="Qwen/Qwen3.6-35B-A3B",
messages=[{"role": "user", "content": "Write a quicksort in Python."}],
max_tokens=4096,
temperature=0.0
)
print(response.choices[0].message.content)
https://huggingface.co/z-lab/Qwen3.6-35B-A3B-DFlash#benchmark-resultsBenchmark Results
**Setup:**Single NVIDIA B200, SGLang, thinking enabled, max output length 4096. We report end-to-end throughput, including prefill time. See ourGitHub repositoryfor reproduction scripts.
https://huggingface.co/z-lab/Qwen3.6-35B-A3B-DFlash#throughput-and-speedupThroughput and Speedup
DFlash achieves up to2.9xspeedup at concurrency 1.
Tokens/sec (speedup vs. autoregressive baseline)
Block Size = 16
TaskConcurrencyARDFlashMath5001234682 (2.9x)812663138 (2.5x)1619544813 (2.5x)3227556520 (2.4x)GSM8K1235556 (2.4x)812362564 (2.1x)1618863821 (2.0x)3226995239 (1.9x)HumanEval1238603 (2.5x)812552800 (2.2x)1619444208 (2.2x)3227675782 (2.1x)MBPP1235559 (2.4x)812242538 (2.1x)1619483816 (2.0x)3227805378 (1.9x)MT-Bench1233442 (1.9x)812382028 (1.6x)1618852997 (1.6x)3226334034 (1.5x)Alpaca1235393 (1.7x)812211782 (1.5x)1618442567 (1.4x)3225793689 (1.4x) Block Size = 8
TaskConcurrencyARDFlashMath5001234617 (2.6x)812662839 (2.2x)1619544465 (2.3x)3227556614 (2.4x)GSM8K1235540 (2.3x)812362466 (2.0x)1618863899 (2.1x)3226995713 (2.1x)HumanEval1238561 (2.4x)812552655 (2.1x)1619444135 (2.1x)3227676059 (2.2x)MBPP1235497 (2.1x)812242324 (1.9x)1619483636 (1.9x)3227804884 (1.8x)MT-Bench1233438 (1.9x)812382060 (1.7x)1618853182 (1.7x)3226334720 (1.8x)Alpaca1235407 (1.7x)812211880 (1.5x)1618442903 (1.6x)3225794115 (1.6x)
https://huggingface.co/z-lab/Qwen3.6-35B-A3B-DFlash#acceptance-lengthAcceptance Length
TaskB8B16Math5005.567.35GSM8K5.216.73HumanEval5.096.44MBPP4.785.83MT-Bench4.205.14Alpaca3.944.62
https://huggingface.co/z-lab/Qwen3.6-35B-A3B-DFlash#acknowledgementsAcknowledgements
Special thanks toDavid Wangfor his outstanding engineering support on this project. We are also grateful toModal,InnoMatrix, andYotta Labsfor providing the compute resources used to train this draft model.
https://huggingface.co/z-lab/Qwen3.6-35B-A3B-DFlash#citationCitation
If you find DFlash useful, please cite our work. To share feedback on DFlash or request new model support, please fill out this form:DFlash Feedback.
@article{chen2026dflash,
title = {{DFlash: Block Diffusion for Flash Speculative Decoding}},
author = {Chen, Jian and Liang, Yesheng and Liu, Zhijian},
journal = {arXiv preprint arXiv:2602.06036},
year = {2026}
}
Similar Articles
z-lab/Qwen3.6-27B-DFlash
This article introduces Qwen3.6-27B-DFlash, a specialized drafter model for DFlash, a novel speculative decoding method using block diffusion to accelerate inference speed. It provides installation instructions for vLLM and SGLang to enable parallel drafting with the target Qwen3.6-27B model.
z-lab/Qwen3.8-27B-DFlash2
Introduces DFlash 2, a block-diffusion drafter for speculative decoding with the Qwen3.8-27B model, demonstrating improved acceptance length and throughput in benchmarks.
incoai/Qwen3.8-27B-DFlash2
DFlash 2 is a block-diffusion draft model for speculative decoding that improves inference speed for the Qwen3.8-27B language model, offering higher acceptance length and throughput in benchmarks.
@charles_irl: Speculation Is All You Need. In this blog post, we announce the co-release (w/ Z Lab) of six more state-of-the-art DFla…
Modal and Z Lab release six new DFlash speculative decoding draft models for Qwen 3.x, achieving over 1000 tokens per second on a B200 and arguing that speculative decoding is the most impactful inference optimization.
z-lab/gemma-4-31B-it-DFlash
Z-lab released DFlash, a speculative decoding drafter model for Gemma-4-31B-it that uses lightweight block diffusion to draft multiple tokens in parallel, achieving up to 5.8x speedup over autoregressive baseline.