@steeve: Initial @Zai_org's DFlash implementation in @zml_ai (and soon in zml/llmd)
Summary
Initial DFlash implementation by Zai_org is integrated into ZML AI, with plans to include it in zml/llmd.
View Cached Full Text
Cached at: 05/15/26, 09:07 PM
Initial @Zai_org’s DFlash implementation in @zml_ai (and soon in zml/llmd) https://t.co/8eFipmepMf
Similar Articles
@zhijianliu_: DFlash is now running in a production inference stack. More draft models coming soon. https://github.com/z-lab/dflash
DFlash is a lightweight block diffusion model for speculative decoding, now running in production with support for various LLMs like Qwen and Gemma.
z-lab/dflash
DFlash introduces a block diffusion method for flash speculative decoding to enhance inference speed in large language models.
@modal: We worked with @lmsysorg and http://z-lab.ai to - integrate DFlash spec into @sgl_project - make it faster with overlap…
Modal collaborated with LMSys and Z Lab to integrate DFlash speculative decoding into SGLang, achieving up to 4.3x throughput improvement over baseline and 1.5x over native multi-token prediction for large language models.
DFlash support merged into llama.cpp
DFlash support has been merged into llama.cpp, improving inference performance for compatible models.
@bstnxbt: dflash-mlx v0.1.6 is out. Biggest agentic update so far: ► much more usable for real OpenCode / coding-agent sessions ►…
dflash-mlx v0.1.6 is released with major agentic improvements, including adaptive verification, custom kernels, prefix cache improvements, and broader compatibility with agentic coding tools like OpenCode, aider, and Continue.