@charles_irl: On Friday, we released six new state-of-the-art drafters for accelerated inference. We also put out a blog post on why …
Summary
On Friday, we released six new state-of-the-art drafters for accelerated inference, along with a blog post on speculative decoding and a roofline model tool to estimate speedups.
View Cached Full Text
Cached at: 06/22/26, 05:32 AM
On Friday, we released six new state-of-the-art drafters for accelerated inference.
We also put out a blog post on why spec dec is so great. Supporting that was a roofline model of speedup from speculation.
Play with it in our LLM Engineer’s Almanac:
https://t.co/udJXMQWlIW https://t.co/Dk7ULxhp54
LLM Engineer’s Almanac - Spec Dec Roofline Model (Speedup ratio)
Source: https://modal.com/llm-almanac/spec-dec-roofline

γ*=16 (max), 1.6x speedup
Sequence length4,096 tok/seq
5122k8k32k131k
Acceptance probability75%
Relative cost per token10%
Acceptance probability89%
Relative cost per block10%
This modeling system usesroofline analysisto estimate the speedups from speculative decoding for different draft lengths applied to different models running on different hardware. It is only a model! It tends to underestimate the benefit whenoverheadis a major contributor to latency, e.g. small batch sizes on small models.
The roofline model used here was inspired by the work ofFergus FinnofDoubleword. In particular, the implementation was derived usinghis DeepSeek-V4 Flash B200 optimal draft length estimatoras a reference.
Similar Articles
@charles_irl: If you're interested in speculative decoding, take some time to grok this chart! And read the article from @haoailab.ht…
A roofline model from the LLM Engineer's Almanac estimates speedups from speculative decoding for different draft lengths across models and hardware, with a note that it may underestimate benefits when overhead is significant.
@jimmysmith1919: Another nice release today. New draft models for speculative decoding of several of our LFM2.5 models. 1.2B: https://hu…
LiquidAI releases draft models for speculative decoding to accelerate their LFM2.5 models, achieving up to 2× faster inference on H100 and Apple silicon without quality degradation.
@charles_irl: Speculation Is All You Need. In this blog post, we announce the co-release (w/ Z Lab) of six more state-of-the-art DFla…
Modal and Z Lab release six new DFlash speculative decoding draft models for Qwen 3.x, achieving over 1000 tokens per second on a B200 and arguing that speculative decoding is the most impactful inference optimization.
@mohitwt_: Day 20/30 of Inference Engineering building a speculative decoding runtime that drafts multiple tokens ahead with a sma…
A developer is building Octane, a speculative decoding runtime for local LLM inference on consumer hardware, aiming for 2-3x speedup with exact output quality. Currently in active development with paged KV cache, continuous batching, and batched attention implemented.
DeepSeek open-sources inference optimizations with 60–85% faster generation [pdf]
DeepSeek open-sourced DeepSpec, a full-stack codebase for training and evaluating draft models for speculative decoding, enabling 60-85% faster generation. It includes data preparation, training, and evaluation scripts with support for multiple draft model algorithms (DSpark, DFlash, Eagle3).