z-lab/Qwen3.6-35B-A3B-DFlash

Hugging Face Models Trending Models

Summary

z-lab releases DFlash, a speculative decoding drafter that uses a lightweight block-diffusion model to draft 15–16 tokens in parallel, yielding up to 2.9× speedup for Qwen3.6-35B-A3B inference.

Task: text-generation Tags: transformers, safetensors, qwen3, feature-extraction, dflash, speculative-decoding, block-diffusion, draft-model, efficiency, qwen, diffusion-language-model, text-generation, custom_code, arxiv:2602.06036, license:mit, text-generation-inference, endpoints_compatible, region:us
Original Article
View Cached Full Text

Cached at: 04/21/26, 07:47 PM

z-lab/Qwen3.6-35B-A3B-DFlash · Hugging Face

Source: https://huggingface.co/z-lab/Qwen3.6-35B-A3B-DFlash Paper|GitHub|Blog

DFlashis a speculative decoding method that uses a lightweightblock diffusionmodel to draft multiple tokens in parallel. This is the drafter model, which must be paired withQwen/Qwen3.6-35B-A3B.

DFlash Architecture

https://huggingface.co/z-lab/Qwen3.6-35B-A3B-DFlash#quick-startQuick Start

https://huggingface.co/z-lab/Qwen3.6-35B-A3B-DFlash#installationInstallation

vLLM:

uv pip install vllm
uv pip install -U vllm --torch-backend=auto --extra-index-url https://wheels.vllm.ai/nightly

SGLang:

uv pip install "git+https://github.com/sgl-project/sglang.git@refs/pull/20547/head#subdirectory=python"

https://huggingface.co/z-lab/Qwen3.6-35B-A3B-DFlash#launch-serverLaunch Server

vLLM:

vllm serve Qwen/Qwen3.6-35B-A3B \
  --speculative-config '{"method": "dflash", "model": "z-lab/Qwen3.6-35B-A3B-DFlash", "num_speculative_tokens": 15}' \
  --attention-backend flash_attn \
  --max-num-batched-tokens 32768

SGLang:

# Optional: enable schedule overlapping (experimental, may not be stable)
# export SGLANG_ENABLE_SPEC_V2=1
# export SGLANG_ENABLE_DFLASH_SPEC_V2=1
# export SGLANG_ENABLE_OVERLAP_PLAN_STREAM=1

python -m sglang.launch_server \
    --model-path Qwen/Qwen3.6-35B-A3B \
    --speculative-algorithm DFLASH \
    --speculative-draft-model-path z-lab/Qwen3.6-35B-A3B-DFlash \
    --speculative-num-draft-tokens 16 \
    --tp-size 1 \
    --attention-backend fa3 \
    --mem-fraction-static 0.75 \
    --mamba-scheduler-strategy extra_buffer \
    --trust-remote-code

**Tip:**For long-context or agentic workloads, add\-\-speculative\-dflash\-draft\-window\-size WINDOW\_SIZEto enable sliding-window attention for the drafter.

https://huggingface.co/z-lab/Qwen3.6-35B-A3B-DFlash#usageUsage

from openai import OpenAI

client = OpenAI(base_url="http://localhost:30000/v1", api_key="EMPTY")

response = client.chat.completions.create(
    model="Qwen/Qwen3.6-35B-A3B",
    messages=[{"role": "user", "content": "Write a quicksort in Python."}],
    max_tokens=4096,
    temperature=0.0
)
print(response.choices[0].message.content)

https://huggingface.co/z-lab/Qwen3.6-35B-A3B-DFlash#benchmark-resultsBenchmark Results

**Setup:**Single NVIDIA B200, SGLang, thinking enabled, max output length 4096. We report end-to-end throughput, including prefill time. See ourGitHub repositoryfor reproduction scripts.

https://huggingface.co/z-lab/Qwen3.6-35B-A3B-DFlash#throughput-and-speedupThroughput and Speedup

DFlash achieves up to2.9xspeedup at concurrency 1.

Tokens/sec (speedup vs. autoregressive baseline)

Block Size = 16

TaskConcurrencyARDFlashMath5001234682 (2.9x)812663138 (2.5x)1619544813 (2.5x)3227556520 (2.4x)GSM8K1235556 (2.4x)812362564 (2.1x)1618863821 (2.0x)3226995239 (1.9x)HumanEval1238603 (2.5x)812552800 (2.2x)1619444208 (2.2x)3227675782 (2.1x)MBPP1235559 (2.4x)812242538 (2.1x)1619483816 (2.0x)3227805378 (1.9x)MT-Bench1233442 (1.9x)812382028 (1.6x)1618852997 (1.6x)3226334034 (1.5x)Alpaca1235393 (1.7x)812211782 (1.5x)1618442567 (1.4x)3225793689 (1.4x) Block Size = 8

TaskConcurrencyARDFlashMath5001234617 (2.6x)812662839 (2.2x)1619544465 (2.3x)3227556614 (2.4x)GSM8K1235540 (2.3x)812362466 (2.0x)1618863899 (2.1x)3226995713 (2.1x)HumanEval1238561 (2.4x)812552655 (2.1x)1619444135 (2.1x)3227676059 (2.2x)MBPP1235497 (2.1x)812242324 (1.9x)1619483636 (1.9x)3227804884 (1.8x)MT-Bench1233438 (1.9x)812382060 (1.7x)1618853182 (1.7x)3226334720 (1.8x)Alpaca1235407 (1.7x)812211880 (1.5x)1618442903 (1.6x)3225794115 (1.6x)

https://huggingface.co/z-lab/Qwen3.6-35B-A3B-DFlash#acceptance-lengthAcceptance Length

TaskB8B16Math5005.567.35GSM8K5.216.73HumanEval5.096.44MBPP4.785.83MT-Bench4.205.14Alpaca3.944.62

https://huggingface.co/z-lab/Qwen3.6-35B-A3B-DFlash#acknowledgementsAcknowledgements

Special thanks toDavid Wangfor his outstanding engineering support on this project. We are also grateful toModal,InnoMatrix, andYotta Labsfor providing the compute resources used to train this draft model.

https://huggingface.co/z-lab/Qwen3.6-35B-A3B-DFlash#citationCitation

If you find DFlash useful, please cite our work. To share feedback on DFlash or request new model support, please fill out this form:DFlash Feedback.

@article{chen2026dflash,
  title   = {{DFlash: Block Diffusion for Flash Speculative Decoding}},
  author  = {Chen, Jian and Liang, Yesheng and Liu, Zhijian},
  journal = {arXiv preprint arXiv:2602.06036},
  year    = {2026}
}

Similar Articles

z-lab/Qwen3.6-27B-DFlash

Hugging Face Models Trending

This article introduces Qwen3.6-27B-DFlash, a specialized drafter model for DFlash, a novel speculative decoding method using block diffusion to accelerate inference speed. It provides installation instructions for vLLM and SGLang to enable parallel drafting with the target Qwen3.6-27B model.

z-lab/Qwen3.8-27B-DFlash2

Hugging Face Models Trending

Introduces DFlash 2, a block-diffusion drafter for speculative decoding with the Qwen3.8-27B model, demonstrating improved acceptance length and throughput in benchmarks.

incoai/Qwen3.8-27B-DFlash2

Hugging Face Models Trending

DFlash 2 is a block-diffusion draft model for speculative decoding that improves inference speed for the Qwen3.8-27B language model, offering higher acceptance length and throughput in benchmarks.

z-lab/gemma-4-31B-it-DFlash

Hugging Face Models Trending

Z-lab released DFlash, a speculative decoding drafter model for Gemma-4-31B-it that uses lightweight block diffusion to draft multiple tokens in parallel, achieving up to 5.8x speedup over autoregressive baseline.