@charles_irl: On Friday, we released six new state-of-the-art drafters for accelerated inference. We also put out a blog post on why …

X AI KOLs Following Models

Summary

On Friday, we released six new state-of-the-art drafters for accelerated inference, along with a blog post on speculative decoding and a roofline model tool to estimate speedups.

On Friday, we released six new state-of-the-art drafters for accelerated inference. We also put out a blog post on why spec dec is so great. Supporting that was a roofline model of speedup from speculation. Play with it in our LLM Engineer's Almanac: https://t.co/udJXMQWlIW https://t.co/Dk7ULxhp54
Original Article
View Cached Full Text

Cached at: 06/22/26, 05:32 AM

On Friday, we released six new state-of-the-art drafters for accelerated inference.

We also put out a blog post on why spec dec is so great. Supporting that was a roofline model of speedup from speculation.

Play with it in our LLM Engineer’s Almanac:

https://t.co/udJXMQWlIW https://t.co/Dk7ULxhp54


LLM Engineer’s Almanac - Spec Dec Roofline Model (Speedup ratio)

Source: https://modal.com/llm-almanac/spec-dec-roofline LLM Engineer’s Almanac

γ*=16 (max), 1.6x speedup

Sequence length4,096 tok/seq

5122k8k32k131k

Acceptance probability75%

Relative cost per token10%

Acceptance probability89%

Relative cost per block10%

This modeling system usesroofline analysisto estimate the speedups from speculative decoding for different draft lengths applied to different models running on different hardware. It is only a model! It tends to underestimate the benefit whenoverheadis a major contributor to latency, e.g. small batch sizes on small models.

The roofline model used here was inspired by the work ofFergus FinnofDoubleword. In particular, the implementation was derived usinghis DeepSeek-V4 Flash B200 optimal draft length estimatoras a reference.

Similar Articles