Vision Is Not Overhead: One-Pass Block Drafting for Lossless Speculative Decoding in Vision-Language Models

arXiv cs.AI Papers

Summary

The paper introduces GLANCE, a one-pass block drafting method for lossless speculative decoding in vision-language models, achieving up to 2.93x faster generation without changing output.

arXiv:2609.00355v1 Announce Type: new Abstract: Speculative decoding accelerates generation without changing its output, yet on vision-language models (VLMs) it has been caught in a self-defeating cycle. The drafter stays autoregressive, so it must stay small. A small drafter cannot afford the image at every step, so vision is compressed, pruned, or hidden. A drafter cut off from the image is then least reliable exactly where the image makes text predictable. We present GLANCE, the first one-pass block drafter that is lossless on an unmodified VLM target, and it breaks the cycle at both ends. A block-diffusion head reads the target's already-fused vision-language state, so vision costs the drafter nothing, and fills a whole block in one forward pass, so depth costs no sequential steps. A wide candidate tree is verified in one target pass, and every audited prompt reproduces greedy decoding exactly. Grounded workloads reward this most, entering a verbatim-copy regime whose long runs cost an autoregressive drafter a pass for every token and a block drafter one in total. Under one engine and one round budget, GLANCE decodes up to 2.93x faster than autoregression, from one draft pass a round where the production EAGLE3-VL head takes eight, and accepts 2.7x longer blocks than an EAGLE-3 head trained on the same corpus. One law organizes these results. Accepted length is set by the target's next-token entropy, with a fitted slope that steepens with grounding across all five tasks. The law transfers across targets and modalities and names its own boundary, since free-running text still favors a chain. Our code is available at https://github.com/js-lee-AI/GLANCE.
Original Article
View Cached Full Text

Cached at: 09/02/26, 06:01 AM

# Vision Is Not Overhead: One-Pass Block Drafting for Lossless Speculative Decoding in Vision-Language Models
Source: [https://arxiv.org/html/2609.00355](https://arxiv.org/html/2609.00355)
Seongtae HongDongyub Jude LeeChanjun ParkJaehyung SeoSugyeong Eo\\correspondingHeuiseok Lim\\corresponding

###### Abstract

Speculative decoding accelerates generation without changing its output, yet on vision\-language models \(VLMs\) it has been caught in a self\-defeating cycle\. The drafter stays autoregressive, so it must stay small\. A small drafter cannot afford the image at every step, so vision is compressed, pruned, or hidden\. A drafter cut off from the image is then least reliable exactly where the image makes text predictable\. We present GLANCE, the first one\-pass block drafter that is lossless on an unmodified VLM target, and it breaks the cycle at both ends\. A block\-diffusion head reads the target’s already\-fused vision\-language state, so vision costs the drafter nothing, and fills a whole block in one forward pass, so depth costs no sequential steps\. A wide candidate tree is verified in one target pass, and every audited prompt reproduces greedy decoding exactly\. Grounded workloads reward this most, entering a verbatim\-copy regime whose long runs cost an autoregressive drafter a pass for every token and a block drafter one in total\. Under one engine and one round budget, GLANCE decodes up to 2\.93x faster than autoregression, from one draft pass a round where the production EAGLE3\-VL head takes eight, and accepts 2\.7x longer blocks than an EAGLE\-3 head trained on the same corpus\. One law organizes these results\. Accepted length is set by the target’s next\-token entropy, with a fitted slope that steepens with grounding across all five tasks\. The law transfers across targets and modalities and names its own boundary, since free\-running text still favors a chain\. Our code is available athttps://github\.com/js\-lee\-AI/GLANCE\.

1Korea University2Zoom Communications3Soongsil University

4Konkuk University5Yonsei University

\{omanma1928, ghdchlwls123, limhseok\}@korea\.ac\.kr, jude\.lee@zoom\.us, chanjun\.park@ssu\.ac\.kr, seojae777@konkuk\.ac\.kr, s\.eo@yonsei\.ac\.kr

## 1Introduction

Vision\-language models \(VLMs\) increasingly serve workloads whose output is a long stretch of text about an image, describing a photograph, reading a scanned document, pulling numbers off a chart\([Qwen Team 2025](https://arxiv.org/html/2609.00355#bib.bib24);[Bai et al\. 2025](https://arxiv.org/html/2609.00355#bib.bib25)\)\. Generation stays autoregressive, so every output token needs a full forward pass of the target, and at small batch sizes that loop is bound by memory bandwidth rather than compute\. The visual encoder runs once, at prefill, while the loop runs on every token, so serving grounded generation faster means attacking the loop rather than the encoder\.

Speculative decoding removes part of that cost without changing the output\. A cheap drafter proposes several future tokens and the target verifies them in one forward pass, committing the longest correct prefix\([Leviathan et al\. 2023](https://arxiv.org/html/2609.00355#bib.bib1);[Chen et al\. 2023](https://arxiv.org/html/2609.00355#bib.bib2)\)\. The modern form drafts from the target’s own hidden features\([Cai et al\. 2024](https://arxiv.org/html/2609.00355#bib.bib7);[Li et al\. 2024b](https://arxiv.org/html/2609.00355#bib.bib3);[Li et al\. 2024a](https://arxiv.org/html/2609.00355#bib.bib4)\)over a tree of candidates\([Miao et al\. 2024](https://arxiv.org/html/2609.00355#bib.bib8);[Chen et al\. 2024b](https://arxiv.org/html/2609.00355#bib.bib34)\), and EAGLE\-3\([Li et al\. 2025b](https://arxiv.org/html/2609.00355#bib.bib5)\)is the production standard whose heads vision\-language stacks now ship\.

![Refer to caption](https://arxiv.org/html/2609.00355v1/fig_teaser_vector.png)Figure 1:Grounded generation is the most draftable\. Top: on a document page the image pins the answer, entropy is low, and one draft pass commits the entire block\. Bottom: open captioning admits many continuations, entropy is high, and the same pass commits only a short prefix\.Work that adapts drafting to VLMs is caught in a cycle of its own premises\. The drafter stays autoregressive, so a candidatekktokens deep costskksequential draft passes, and the drafter must stay small for those passes to be cheap\([AQ\-MedAI 2025](https://arxiv.org/html/2609.00355#bib.bib6)\)\. A small drafter cannot afford the image at every step, so image tokens are compressed, pruned, or hidden from it\([Gagrani et al\. 2024](https://arxiv.org/html/2609.00355#bib.bib23);[Kang et al\. 2025](https://arxiv.org/html/2609.00355#bib.bib15);[Huang et al\. 2025](https://arxiv.org/html/2609.00355#bib.bib16);[Xie et al\. 2025](https://arxiv.org/html/2609.00355#bib.bib17)\)\. Yet depth is what a drafter has to sell, and a drafter cut off from the image is least reliable about exactly the text the image already fixes, so each premise enforces the next, and the cycle prices out the drafts that would pay the most\.

We ask the opposite question\. On grounded workloads, is vision a cost to hide from the drafter, or the reason drafting works at all? When a VLM answers a question over a document, much of what it generates already exists in the image, so the next token is often near\-deterministic, frequently an exact copy off the page, as Figure[1](https://arxiv.org/html/2609.00355#S1.F1)shows\. Predictability is what a drafter converts into speed, and we make it quantitative\. Accepted block length is governed by the target’s next\-token entropy through one law that fits every task we test, and grounded tasks sit at the low\-entropy end, so reading documents and charts accepts the longest blocks and open description the shortest\. A mismatched\-image probe shows the image acts through that entropy channel rather than beside it, and swapping the drafter’s fused visual state for a text\-only LLM’s costs the most acceptance exactly where drafting is wide and deep\.

![Refer to caption](https://arxiv.org/html/2609.00355v1/fig_framework_vector.png)Figure 2:One\-pass block speculative decoding for VLMs\. Top: the frozen Qwen3\-VL\-8B target and its input\. The block head reads the*fused*vision\-language hidden states at five kept layers, never the raw visual tokens\. Bottom, one decoding round left to right: one draft pass fills the whole block at every offset at once, the highest\-scoring prefixes form a prefix\-closed candidate tree, and a single ancestor\-masked target pass verifies all of them, committing the longest path the target itself would have taken\.Low entropy pays off only when long, wide candidate sets can be proposed cheaply, so we draft in one pass\. We present GLANCE, for grounded block drafting with one\-pass candidate expansion, and it exits the cycle at both ends, inheriting vision and filling depth at once\. A block\-diffusion head\([Chen et al\. 2026](https://arxiv.org/html/2609.00355#bib.bib9);[Arriola et al\. 2025](https://arxiv.org/html/2609.00355#bib.bib12)\)reads the frozen target’s own fused vision\-language state and draws a distribution over every position of a whole block from a single draft pass\. The highest\-ranked prefixes form a wide candidate tree, and one target pass verifies all of them\. We prove the committed path is exactly the target’s own greedy decoding, and we gate every run on that equality rather than assume it\. Because one draft pass has already paid for the whole block, width becomes the cheap axis, bought inside the verify rather than with more draft passes\([Ringel and Romano 2026](https://arxiv.org/html/2609.00355#bib.bib10);[Zhang et al\. 2026b](https://arxiv.org/html/2609.00355#bib.bib11)\)\.

The result is a drafter that turns grounding into speed\. On document, infographic and chart question answering GLANCE decodes up to2\.93×2\.93\\timesfaster than autoregression, from one draft pass a round where the production vision\-language head takes eight, and every committed token is the target’s own\. Trained on the same corpus and schedule as an EAGLE\-3 head, it accepts2\.7×2\.7\\timeslonger blocks\. The law behind the gains transfers to a second vision\-language target and to speech, code and chat\.

Our contributions are:

- •We rediagnose why drafting has underperformed on VLMs\. An autoregressive drafter and stripped\-down vision access enforce each other, and together they discard the workload’s best asset, grounded generation’s long verbatim runs, which only a one\-pass drafter harvests whole\.
- •We build the first one\-pass block drafter that is lossless on an unmodified vision\-language target, combining a block\-diffusion head over fused vision\-language states, wide\-tree verification in a single target pass, and measured bitwise equality with greedy decoding\.
- •We make the mechanism quantitative\. Accepted length is governed by the target’s next\-token entropy, a mismatched\-image probe shows the image acts through that channel, and at the grounded floor a verbatim\-copy analysis lifts measured acceptance above the law’s own cap on every grounded task\. OCR\-targeted retraining lifts acceptance where the law predicts headroom\.
- •We settle the deployment claim inside one production engine on five task families, and map the law’s scope across drafters, targets and modalities\.

## 2Related Work

Draft heads and tree verification\.Speculative decoding uses a rejection\-sampling accept rule that provably preserves the target distribution\([Leviathan et al\. 2023](https://arxiv.org/html/2609.00355#bib.bib1);[Chen et al\. 2023](https://arxiv.org/html/2609.00355#bib.bib2)\)\. The practical line drafts from the target’s own hidden features with a lightweight head: independent offset\-wise heads\([Cai et al\. 2024](https://arxiv.org/html/2609.00355#bib.bib7)\), sequentially dependent heads\([Ankner et al\. 2024](https://arxiv.org/html/2609.00355#bib.bib38)\), and feature\-level autoregression with dynamic candidate trees\([Li et al\. 2024b](https://arxiv.org/html/2609.00355#bib.bib3);[Li et al\. 2024a](https://arxiv.org/html/2609.00355#bib.bib4)\), which EAGLE\-3 extends with training\-time test and multi\-layer feature fusion\([Li et al\. 2025b](https://arxiv.org/html/2609.00355#bib.bib5)\)\. Many candidates are verified at once by packing a prefix tree into a single ancestor\-masked target pass\([Miao et al\. 2024](https://arxiv.org/html/2609.00355#bib.bib8)\), with later work on tree shape\([Chen et al\. 2024b](https://arxiv.org/html/2609.00355#bib.bib34);[Wang et al\. 2025a](https://arxiv.org/html/2609.00355#bib.bib35)\)and on theory and benchmarking\([Yin et al\. 2024](https://arxiv.org/html/2609.00355#bib.bib42);[Huang et al\. 2024](https://arxiv.org/html/2609.00355#bib.bib45);[Xia et al\. 2024](https://arxiv.org/html/2609.00355#bib.bib40)\)\.

Speculative decoding for VLMs\.The first study of speculative decoding for multimodal LLMs found a text\-only drafter a strong baseline\([Gagrani et al\. 2024](https://arxiv.org/html/2609.00355#bib.bib23)\)\. Subsequent work takes one of two routes\. One shrinks what the drafter must read, by compressing image tokens into adaptor embeddings\([Kang et al\. 2025](https://arxiv.org/html/2609.00355#bib.bib15)\), pruning visual tokens\([Huang et al\. 2025](https://arxiv.org/html/2609.00355#bib.bib16)\), or hiding them outright\([Xie et al\. 2025](https://arxiv.org/html/2609.00355#bib.bib17)\)\. The other gives a small drafter its own cheap vision path through distillation and target\-feature fusion\([Ganesan et al\. 2025](https://arxiv.org/html/2609.00355#bib.bib18);[Hu et al\. 2025](https://arxiv.org/html/2609.00355#bib.bib19)\)\. Semi\-autoregressive multimodal drafting relaxes one token per pass without an exactness guarantee\([Wang et al\. 2025b](https://arxiv.org/html/2609.00355#bib.bib21)\), and the setting has a benchmark\([Shen et al\. 2026](https://arxiv.org/html/2609.00355#bib.bib22)\)\. Both routes keep the drafter autoregressive and engineer the drafter’s vision access rather than inheriting it\. We inherit it in one pass\.

Block drafting and diffusion drafters\.A parallel line removes the sequential bottleneck inside the drafter itself: lookahead embeddings\([Monea et al\. 2023](https://arxiv.org/html/2609.00355#bib.bib36)\), Jacobi decoding\([Fu et al\. 2024](https://arxiv.org/html/2609.00355#bib.bib37)\), block\-parallel discrete\-diffusion language models\([Nie et al\. 2025](https://arxiv.org/html/2609.00355#bib.bib13);[Arriola et al\. 2025](https://arxiv.org/html/2609.00355#bib.bib12)\), diffusion drafters with an autoregressive verifier\([Christopher et al\. 2025](https://arxiv.org/html/2609.00355#bib.bib39);[Li et al\. 2025a](https://arxiv.org/html/2609.00355#bib.bib14);[Cheng et al\. 2025](https://arxiv.org/html/2609.00355#bib.bib44)\), and a block\-diffusion head on the frozen target’s hidden state\([Chen et al\. 2026](https://arxiv.org/html/2609.00355#bib.bib9)\)\. One\-pass drafting changes the economics of tree width, settled for text\-only block drafting\([Ringel and Romano 2026](https://arxiv.org/html/2609.00355#bib.bib10);[Zhang et al\. 2026b](https://arxiv.org/html/2609.00355#bib.bib11)\)and reused here unchanged\. Block drafting has reached vision\-language targets only by converting the target itself into a self\-speculating diffusion model\([Wu et al\. 2026](https://arxiv.org/html/2609.00355#bib.bib33);[Zhang et al\. 2026a](https://arxiv.org/html/2609.00355#bib.bib43)\), which forfeits the original model’s output\. We keep the target frozen and its output exact\.

## 3Preliminaries

### 3\.1Speculative Decoding and Acceptance

Letp\(⋅∣x\)p\(\\cdot\\mid x\)be the frozen target VLM’s next\-token distribution given the multimodal prefixxx\(text tokens together with the encoded image\)\. Autoregressive greedy decoding emitsarg​maxw⁡p​\(w∣x\)\\argmax\_\{w\}p\(w\\mid x\), one token for each target forward\. Speculative decoding instead proceeds in*rounds*\([Leviathan et al\. 2023](https://arxiv.org/html/2609.00355#bib.bib1);[Chen et al\. 2023](https://arxiv.org/html/2609.00355#bib.bib2)\)\. Given the committed prefix and a pending*root*tokenbb, a drafter proposes candidate continuations and the target commits the longest prefix consistent with itself\. At temperature00that accept rule is exact prefix matching against the target’s own greedy continuationY=\(Y1,Y2,…\)Y=\(Y\_\{1\},Y\_\{2\},\\dots\)afterbb, whereYk=arg​maxwp\(w∣x∘b∘Y1:k−1\)Y\_\{k\}=\\argmax\_\{w\}p\(w\\mid x\\circ b\\circ Y\_\{1:k\-1\}\)\. Letaabe the number of accepted nonroot tokens in a round\. Each round commitsa\+1a\+1tokens, the accepted prefix plus one token read off the last verified distribution\.

The central quantity of this paper is the*average acceptance length*, the number of tokens committed in one verification round,

τ=𝔼⁡\[a\+1\]=𝔼⁡\[a\]\+1,\\tau\\;=\\;\\mathbb\{E\}\[a\+1\]\\;=\\;\\mathbb\{E\}\[a\]\+1,\(1\)so plain autoregressive decoding hasτ=1\\tau\{=\}1\. Systems that instead log the bonus\-excluded count𝔼⁡\[a\]\\mathbb\{E\}\[a\]\([Kang et al\. 2025](https://arxiv.org/html/2609.00355#bib.bib15)\)are raised by the committed token so every row sits on one scale\. A higherτ\\taumeans fewer target passes for each emitted token, and the speedup isτ\\taudiscounted by drafting and verification overhead\.

### 3\.2One\-Pass Block Drafting

Our drafter follows the block\-diffusion interface\([Chen et al\. 2026](https://arxiv.org/html/2609.00355#bib.bib9);[Ringel and Romano 2026](https://arxiv.org/html/2609.00355#bib.bib10)\), stated as an assumption since everything downstream depends only on it\.

###### Assumption 1\(One\-pass marginal ranking interface\)\.

Conditioned on the cached prefixxxand the pending rootbb, the drafter returns one marginalqj\(⋅∣x,b\)q\_\{j\}\(\\cdot\\mid x,b\)at every offsetj=1,…,Lj=1,\\dots,Lof the block in a single forward pass, before any within\-block token is committed, and a candidate prefix is ranked by its plug\-in scoreπ^\(y1:ℓ\)=∏j≤ℓqj\(yj∣x,b\)\\hat\{\\pi\}\(y\_\{1:\\ell\}\)=\\prod\_\{j\\leq\\ell\}q\_\{j\}\(y\_\{j\}\\mid x,b\)\.

###### Definition 1\(Candidate tree and accept rule\)\.

A*candidate tree*SSis a prefix\-closed set of nonempty strings after the rootbb, of depth at mostLLand*budget*N=\|S\|N=\|S\|\. Packed into one ancestor\-masked target pass, the row of a depth\-ddnodey1:dy\_\{1:d\}yields exactlyp\(⋅∣x∘b∘y1:d\)p\(\\cdot\\mid x\\circ b\\circ y\_\{1:d\}\)\. WithYYthe target’s greedy continuation, the*accepted length*isA\(S\)=max\{k≥0:Y1:j∈Sfor allj≤k\}A\(S\)=\\max\\\{k\\geq 0:Y\_\{1:j\}\\in S\\ \\text\{for all \}j\\leq k\\\}\. The walk commitsY1:A⁡\(S\)Y\_\{1:A\(S\)\}and emits the next target\-greedy token from the last verified row as the new root\.

Drafting therefore costs one pass for the block,*independent of tree width*\. GrowingNNadds verifier FLOPs inside that single pass but no sequential draft steps\. An autoregressive drafter instead pays depth in sequential passes, five a round for the deployed EAGLE3\-VL default and eight where that head runs fastest\.

## 4Method

We call the resulting decoderGLANCE\. Figure[2](https://arxiv.org/html/2609.00355#S1.F2)draws a decoding round in two steps\. First, a block head drafts a whole block of future tokens in one forward pass \(stage33\)\. Then the target verifies a wide tree of those candidates in one forward pass and commits its own greedy prefix \(stages44and55\)\. The head is trained and the target is frozen\.

### 4\.1Drafting on Fused Vision\-Language States

The drafter is a55\-layer block head in the DFlash family\([Chen et al\. 2026](https://arxiv.org/html/2609.00355#bib.bib9)\)with block sizeB=16B\{=\}16and1\.051\.05B parameters, drawn in stage33of Figure[2](https://arxiv.org/html/2609.00355#S1.F2), attached to a frozen Qwen3\-VL\-8B target\. It never reads raw visual tokens\. Its cross\-attention keys and values come from the target’s fused hidden states at five kept layers\{1,9,17,25,33\}\\\{1,9,17,25,33\\\}, following EAGLE\-3’s multi\-layer feature fusion\([Li et al\. 2025b](https://arxiv.org/html/2609.00355#bib.bib5)\)\. By the time these states exist the target has already merged the image into its text representation, so the drafter inherits the visual grounding at no extra encoding cost\. Given the cached prefix and the pending root token, one forward pass fills the masked block and emits the offset\-wise marginalsq1,…,q15q\_\{1\},\\dots,q\_\{15\}of Assumption[1](https://arxiv.org/html/2609.00355#Thmassumption1)\.

### 4\.2Wide\-Tree Verification and Losslessness

The marginals feed the candidate tree of stage44\. The builder keeps the prefixes of highest plug\-in scoreπ^\\hat\{\\pi\}\(Assumption[1](https://arxiv.org/html/2609.00355#Thmassumption1)\), which forms a prefix\-closed tree\. The tree is packed into a single ancestor\-masked target pass\([Miao et al\. 2024](https://arxiv.org/html/2609.00355#bib.bib8)\)\. Each node’s row reproduces the exact autoregressive conditional at that node \(Definition[1](https://arxiv.org/html/2609.00355#Thmdefinition1)\)\. The decoder then commits the longest target\-greedy path, plus one bonus token read off the last verified row\.

Since drafting already paid for the whole block, the builder reuses the published tree\-economics machinery unchanged\([Ringel and Romano 2026](https://arxiv.org/html/2609.00355#bib.bib10);[Zhang et al\. 2026b](https://arxiv.org/html/2609.00355#bib.bib11)\)\. The budget is a small fixed constant, and Section[5\.4](https://arxiv.org/html/2609.00355#S5.SS4)shows returns flatten once the width is moderate\.

###### Theorem 1\(Lossless greedy equivalence, informal\)\.

Assume the target’s argmax is unique at every visited prefix and its top\-22logit gap there exceeds the packed\-versus\-unpacked logit perturbation\. Then for any candidate tree the greedy walk of Definition[1](https://arxiv.org/html/2609.00355#Thmdefinition1)commits exactly the frozen target’s greedy autoregressive sequence, and in fp32 arithmetic the equality is bitwise\. Full statement and proof in Technical Appendix[B](https://arxiv.org/html/2609.00355#A2)\.

The guarantee is a property of the verifier, not the drafter\. A weak head can only shorten accepted prefixes, never change the output\. Production bf16 kernels can reorder exact ties, so a tie\-aware gate defers to the target’s own token there, as Technical Appendix[B](https://arxiv.org/html/2609.00355#A2)details\.

### 4\.3Training

Only the block head trains, and the vision tower and target decoder stay frozen\. The training data are the target’s own greedy generations on80008000COCO\-Caption2017 and80008000TextVQA prompts,1616K teacher\-forced rows in total\. The objective is offset\-wise top\-11coverage of the target’s own token at each block offset, with a reveal trick that exposes a random prefix so the head learns every reveal length\. The head is warm\-started from the released text\-only Qwen3\-8B DFlash block head and trained for one epoch on a single GPU, with the full training card in Technical Appendix[D](https://arxiv.org/html/2609.00355#A4)\. The recipe contains neither ChartQA nor DocVQA data, yet those end up the two most draftable tasks, the phenomenon Section[6](https://arxiv.org/html/2609.00355#S6)explains\.

## 5Experiments

### 5\.1Setup

#### Models and tasks\.

The primary target is Qwen3\-VL\-8B\-Instruct\([Qwen Team 2025](https://arxiv.org/html/2609.00355#bib.bib24)\), kept frozen\. Qwen3\-VL\-4B and Llama\-3\.2\-11B\-Vision enter only the transfer study\. Tasks span five families ordered from open\-ended to grounded: COCO captioning\([Lin et al\. 2014](https://arxiv.org/html/2609.00355#bib.bib27)\), TextVQA\([Singh et al\. 2019](https://arxiv.org/html/2609.00355#bib.bib26)\), InfographicVQA\([Mathew et al\. 2022](https://arxiv.org/html/2609.00355#bib.bib46)\), DocVQA\([Mathew et al\. 2021](https://arxiv.org/html/2609.00355#bib.bib28)\), and ChartQA\([Masry et al\. 2022](https://arxiv.org/html/2609.00355#bib.bib29)\)\. The rest of the standard suite is multiple\-choice or single\-word, leaving no decode loop to amortize, and ScienceQA stands in for that regime\.

#### Baselines\.

Following the EAGLE series\([Li et al\. 2024b](https://arxiv.org/html/2609.00355#bib.bib3);[Li et al\. 2024a](https://arxiv.org/html/2609.00355#bib.bib4);[Li et al\. 2025b](https://arxiv.org/html/2609.00355#bib.bib5)\)we take the members of its suite that preserve the target’s outputs: the n\-gram lookup floor, classic two\-model speculation at three draft sizes, the production EAGLE3\-VL head\([AQ\-MedAI 2025](https://arxiv.org/html/2609.00355#bib.bib6)\), and ViSpec\([Kang et al\. 2025](https://arxiv.org/html/2609.00355#bib.bib15)\)with its Medusa baseline\. ViSpec carries EAGLE\-2’s dynamic\-tree drafting into the VLM setting\. We train the last two ourselves, since no published VLM drafter releases a head for this target\.

#### Metrics and protocol\.

Every table below inherits one protocol, and a caption states only what departs from it\. Decoding is batch one, decode\-only, capped at256256new tokens, and greedy unless a temperature is marked\. Acceptance lengthτ\\tauof Equation \([1](https://arxiv.org/html/2609.00355#S3.E1)\) counts tokens, not time, so it does not depend on the engine and compares systems directly\. Speedup is the ratio of the mean decode time of one token, autoregressive over speculative, each system against its own baseline on one device\. The n\-gram and EAGLE3\-VL rows of Table[3](https://arxiv.org/html/2609.00355#S5.T3)are the serving engine’s end\-to\-end times, which pull a ratio toward1\.01\.0\. Table[1](https://arxiv.org/html/2609.00355#S5.T1)holds the round budget at3232draft tokens for both drafters, while every other table runs each system at its own operating point, ours a6363\-node tree, soτ\\tauis not one number across tables\. Every greedy run is gated at fp32 on reproducing the target’s own decoding exactly, an equivalence prior VLM work asserts but does not measure\. Prompt counts and the remaining details are in Technical Appendix[D](https://arxiv.org/html/2609.00355#A4)\.

### 5\.2Head\-to\-Head in a Production Engine

Table 1:GLANCE against the production EAGLE3\-VL head with engine, card and round budget held fixed\. The three arms of a task run back to back on one card, both speculative arms CUDA\-graph captured\. Both drafters verify3232draft tokens a round and differ only in how those tokens are produced\.*GLANCE faster by*is the difference in mean decode time\.Both drafters run inside SGLang0\.5\.60\.5\.6on one card, so with engine, card and round budget fixed, the only structural difference left is that EAGLE3\-VL produces its3232draft tokens with eight sequential passes and GLANCE with one\. On the three tasks whose answer is anchored in the image, GLANCE decodes up to2\.93×2\.93\\timesfaster than autoregression and up to7\.6%7\.6\\%faster than the production head, every paired bootstrap interval excluding zero\.

Where the output is free\-running text the eight\-pass chain leads instead, and the split is sharp rather than noisy, at one prompt in101101on captioning and on TextVQA against at least7070on the grounded tasks\. Assumption[1](https://arxiv.org/html/2609.00355#Thmassumption1)names the mechanism, since our head ranks candidates by a product of offset\-wise marginals, exact only where the block’s tokens are conditionally independent given the image\. Near\-determinism delivers that and free\-running text does not, so our deep candidates drift where a chain stays coherent by construction\. Our accepted length is ordered by grounding across all five tasks,2\.91<3\.32<3\.68<3\.93<4\.622\.91<3\.32<3\.68<3\.93<4\.62, while the production head’s is not, placing TextVQA above both DocVQA and InfographicVQA\. And the two tasks we lose are the only two in the training set of Section[4\.3](https://arxiv.org/html/2609.00355#S4.SS3), while all three we win are unseen, so the wins are not a data\-selection effect\.

### 5\.3Acceptance Against Deployed Systems

Matched training\.Both drafter architectures are trained from scratch on one corpus with the same frozen target, global batch, epochs, and framework\([SGLang Team 2026](https://arxiv.org/html/2609.00355#bib.bib41)\), then scored on the same held\-out prompts\. The matched\-training block of Table[2](https://arxiv.org/html/2609.00355#S5.T2)reports the outcome\. GLANCE accepts2\.7×2\.7\\timeslonger blocks pooled over the five domains and4\.0×4\.0\\timeson ChartQA\. The gap replicates on a second training corpus,2\.042\.04against1\.291\.29, and the block’s‡\\ddaggerspeedups show it carrying through to wall\-clock\. The probe isolates training data, schedule, and framework, not the draft budget, which panel \(b\) of Figure[3](https://arxiv.org/html/2609.00355#S5.F3)controls instead with the identical head as a width\-11chain\.

Table 2:Greedy decoding against everything shipped for these targets\. A speedup is a within\-system ratio against that system’s own autoregressive baseline, andτ\\tauis engine\-independent\. The matched\-training block is a held\-out probe, its‡\\ddaggerspeedups timed in one HF harness and read only against each other\. Params is the drafter’s own weights, excluding the frozen target and any bundled copies of its embeddings\. Lossless:✓\\checkmarkexact,×\\timesrelaxed,†\\daggeraudited bitwise identical to greedy decoding\.Deployed systems\.Table[2](https://arxiv.org/html/2609.00355#S5.T2)compares every system shipped for these targets on the metric that survives a change of engine\. Acceptance rises with grounding, and Section[5\.5](https://arxiv.org/html/2609.00355#S5.SS5)traces the gain to the wide tree rather than to the head\. Losslessness we gate rather than assume\. GLANCE reproduces greedy decoding on every audited prompt, while ViSpec and Medusa reproduce none under any tree setting, placing their divergence in the relaxed acceptance rule\. The audit is read at fp32 because in bf16 the target reproduces itself on fewer than half the prompts with no drafter at all, so bf16 disagreement measures the arithmetic and not the drafter\.

Table 3:Sampling atT=1T\{=\}1on the same target, on the three tasks every sampled system shares\. ViSpec and Medusa run greedy only and have no sampling row\. The greedy ordering survives the change of temperature\.ViSpec\.We place the strongest published VLM drafter in both available frames\. On its own Qwen2\.5\-VL\-7B target GLANCE matches or exceeds its accepted length on every task, and is the one system gated for exact reproduction\. On our target, with its head trained by the authors’ official pipeline, no output in any tree configuration matches the target’s greedy text, since the typical\-acceptance rule trades fidelity for length, so we report length\-controlled speedups\.

Figure 3:Systems results on Qwen3\-VL\-8B\. \(a\) Accepted length against*measured*next\-token entropy, by decile, for the three main\-run tasks\. The curves separate by grounding, and DocVQA’s near\-certain end clears the pooled ceiling \(Theorem[2](https://arxiv.org/html/2609.00355#Thmtheorem2)\)\. \(b\) The wide tree against the identical head run as a width\-11chain\. \(c\) Margin against the production head at two context lengths\. \(d\) Speedup against the verifier budgetNNon five tasks\.
### 5\.4Wall\-Clock and Generality

Pushing the same matched pair onto longer prompts, full\-resolution InfographicVQA at two context lengths, locates where the margin lives\. At a22K context GLANCE is a paired\+5\.8%\+5\.8\\%faster than the production head,95%95\\%CI\[\+2\.2,\+9\.4\]\[\+2\.2,\+9\.4\], and it gets there while matching that head’s accepted length from*one*draft forward pass a round against its eight\. Graph capture is what converts the saved passes into saved time, a one\-pass draft being launch\-bound, and with graphs off the cell is parity\. At∼6\.7\{\\sim\}6\.7K the two are level, our accepted length having fallen further than theirs and the round\-cost advantage absorbing the difference\. Technical Appendix[E](https://arxiv.org/html/2609.00355#A5)lists every comparison with its sample size\.

Classic two\-model speculation leads acceptance on four of the five tasks in Table[2](https://arxiv.org/html/2609.00355#S5.T2)yet is a net slowdown on all five, because wall\-clock is decided by what each accepted token costs to draft, not by how much a drafter accepts\. And speedup saturates past a moderate budget on every task, drawn in panel \(d\) of Figure[3](https://arxiv.org/html/2609.00355#S5.F3)and tabulated in Technical Appendix[E](https://arxiv.org/html/2609.00355#A5)\. Width pays most exactly where we win\.

Two regimes bound the method\. Speculation pays only where there is a decode loop to amortize\. On ScienceQA\([Lu et al\. 2022](https://arxiv.org/html/2609.00355#bib.bib30)\), whose answer is a single letter, our loop falls below parity, so this is a method for generation, not classification\. Under sampling, reported in Table[3](https://arxiv.org/html/2609.00355#S5.T3), the committed token is drawn from the target’s own verified row, so the distribution is the target’s by construction and byte\-identity is undefined\.

### 5\.5Ablations

#### What the drafter reads\.

Zeroing the fused conditioning state collapses acceptance to1\.111\.11, so the head genuinely drafts from the target’s context\. Swapping those states for a text\-only LLM’s costs almost nothing on captioning, which recovers97%97\\%of full acceptance, and more as the task becomes grounded, to a minimum of80%80\\%on DocVQA\. The same swap takes a larger share of acceptance from the budget\-3131tree than from the identical width\-11chain on all five tasks, so losing vision costs most exactly when drafting wide and deep\.

#### The round budget, not head size, is the source\.

GLANCE’s head carries2\.6×2\.6\\timesthe parameters of EAGLE3\-VL’s, yet as a width\-11chain the same head gives up at least1\.45×1\.45\\timesof its accepted length on every task, panel \(b\) of Figure[3](https://arxiv.org/html/2609.00355#S5.F3)\. In the forward\-level harness, where the two arms are the same head, the tree decodes1\.36×1\.36\\timesthe chain\. The tree and not the head earns the margin\. Across Table[1](https://arxiv.org/html/2609.00355#S5.T1)the two systems spend within4%4\\%of the same time in a verify round, so the larger head run once costs no more than the small head run eight times, and widening our tree4\.2×4\.2\\timesmoves a round by only a few percent\.

Figure 4:Task\-level law fits inside their two brackets, for all five tasks, ChartQA under its templated prompts\. Axes: root entropyHHin nats against expected accepted length𝔼⁡\[a∣H\]\\mathbb\{E\}\[a\\mid H\], log scale\. The fit of Equation \([2](https://arxiv.org/html/2609.00355#S6.E2)\) stays above the AdaEDL lower bound and below the2​log⁡N/H2\\log N/Hverification ceiling atN=31N\{=\}31\. Points: measured entropy deciles\.
#### Targeted data lifts grounded tasks selectively\.

Retraining the head on OCR\-heavy data lifts DocVQA at both verifier budgets while captioning stays flat, so the gain concentrates where the law predicts headroom\.

## 6Why Grounded Is More Draftable

Every acceptance result above is organized by one scalar: the target’s next\-token entropyHHat the round root\. Entropy has gated drafting before\([Agrawal et al\. 2024](https://arxiv.org/html/2609.00355#bib.bib31);[Tong et al\. 2026](https://arxiv.org/html/2609.00355#bib.bib20);[Mahmoud 2026](https://arxiv.org/html/2609.00355#bib.bib32)\)\. We give it a closed form and test it\. Modeling offset\-wise matches as a survival process with a slowly varying hazard, under two approximations of Technical Appendix[B](https://arxiv.org/html/2609.00355#A2), gives a truncated\-geometric mean, and an affine model of the match logit inHHcloses the form:

logit⁡pm​\(H\)=b0−b1​H,𝔼⁡\[a∣H\]≈pm1−pm,\\operatorname\{logit\}p\_\{\\mathrm\{m\}\}\(H\)=b\_\{0\}\-b\_\{1\}H,\\qquad\\mathbb\{E\}\[a\\mid H\]\\;\\approx\\;\\frac\{p\_\{\\mathrm\{m\}\}\}\{1\-p\_\{\\mathrm\{m\}\}\},\(2\)truncated at the block horizon\. Herelogit⁡u=ln⁡\(u/\(1−u\)\)\\operatorname\{logit\}u=\\ln\(u/\(1\-u\)\),pm​\(H\)∈\(0,1\)p\_\{\\mathrm\{m\}\}\(H\)\\in\(0,1\)is the probability that the drafter’s proposal at a round root of entropyHHmatches the target’s greedy token, andb0,b1\>0b\_\{0\},b\_\{1\}\>0are the zero\-entropy intercept and the entropy slope, fitted for each task on its round log \(values in Technical Appendix[F](https://arxiv.org/html/2609.00355#A6)\)\.

A Fano argument grounds the monotone part, and the affine logit is a surrogate under a temperature family\. The law fits every task with a positive slope that steepens with grounding\. Figure[3](https://arxiv.org/html/2609.00355#S5.F3)\(a\) draws the decile curves for the three main\-run tasks, and Figure[4](https://arxiv.org/html/2609.00355#S5.F4)places all five fits inside two brackets\. Fitted over77517751rounds the pooled law has interceptb0=0\.96b\_\{0\}\{=\}0\.96and slopeb1=0\.63b\_\{1\}\{=\}0\.63\. The per\-task slopes climb monotonically with grounding,0\.460\.46on captioning,0\.580\.58on TextVQA,0\.800\.80on InfoVQA,0\.950\.95on DocVQA and, under ChartQA’s templated prompts,1\.711\.71\. Acceptance has large round\-level scatter, so fits are read against the conditional mean, where the pooledR2R^\{2\}is0\.710\.71, and Technical Appendix[B](https://arxiv.org/html/2609.00355#A2)bounds what any entropy\-measurable predictor can do against raw rounds, a ceiling our fit realizes up to89%89\\%of\.

#### The image acts through entropy\.

Correlation does not show the*image*is responsible, so we intervene on it alone\. The same prompts are decoded with the correct and with a mismatched image on a shared teacher\-forced trajectory\. The correct image lengthens accepted blocks in every task and the effect scales with grounding,\+3\.5%\+3\.5\\%on captioning,\+6\.8%\+6\.8\\%on TextVQA and\+11\.7%\+11\.7\\%on DocVQA, every interval excluding zero\. It also collapses the target’s mean next\-token entropy,0\.260\.26against0\.720\.72nats on DocVQA, and once entropy is held fixed the residual image effect is no longer positive \(pooled coefficient−0\.07\-0\.07\)\. With the feature\-access ablation above, this places the image’s contribution in the entropy channel\.

#### The form transfers\.

If the law is a property of the target rather than of our head, its form should survive a change of drafter, target and modality, and it does on all three axes, with a positive slope every time\. Two\-model drafts on two further vision\-language targets fit atb1=0\.67b\_\{1\}\{=\}0\.67and0\.400\.40, and outside vision the slope stays positive on speech recognition \(0\.490\.49\), chat \(0\.640\.64\), a non\-Qwen backbone \(0\.840\.84\) and code \(1\.761\.76\)\. Grounding is simply one way a task collapses the target’s entropy, and the full fits are in Technical Appendix[F](https://arxiv.org/html/2609.00355#A6)\.

#### The ordering is architecture\-scoped\.

Trained on Llama\-3\.2\-11B\-Vision, where vision enters by cross\-attention rather than as a token prefix, the*method*still works and stays lossless\. The*characterization*does not, since the ordering puts captioning first and the law inverts sign on the grounded short\-answer domains\.

#### The law bends at its own tail\.

At the entropy floor the single\-hazard family gives out, in a direction that favors grounding\.

###### Theorem 2\(Grounded tail, informal\)\.

Equation \([2](https://arxiv.org/html/2609.00355#S6.E2)\) caps the near\-certain conditional mean ateb0e^\{b\_\{0\}\}, yet a verbatim\-copy fraction above an explicit threshold lifts theH→0H\\to 0mean strictly above that cap, and the measured grounded floor sits above it,5\.235\.23against3\.413\.41on DocVQA\. Statement and test in Technical Appendix[B](https://arxiv.org/html/2609.00355#A2)\.

This copy regime is where one\-pass drafting collects its rent, since a run of lengthℓ\\ellcosts an autoregressive drafterℓ\\ellsequential passes and this drafter one\. The grading extends across tasks, since the fraction of rounds accepting eight or more tokens rises with grounding, from0\.0%0\.0\\%on captioning to11\.5%11\.5\\%on ChartQA under its templated prompts, and every grounded task’s floor clears its own cap\.

## 7Conclusion

We presented a lossless one\-pass block drafter for a frozen vision\-language model\. A block\-diffusion head proposes a whole block in one pass from the target’s fused states, and a wide tree turns that parallelism into accepted length, bitwise identical to greedy autoregression\. Our results trace the gains to an exit from the cycle that held VLM drafting back\. They show vision serving as inherited context rather than overhead, depth filled in one pass rather than paid in sequential steps, and grounded generation’s verbatim runs harvested whole\. An entropy law fitted on every task captures when this pays, persists when the drafter, the target and the modality change, and indicates where the method stops, with free\-running text staying a chain’s regime\.

## References

- Agrawalet al\.\(2024\)S\. Agrawal, W\. Jeon, and M\. LeeAdaEDL: early draft stopping for speculative decoding of large language models via an entropy\-based lower bound on token acceptance probability\.External Links:2410\.18351,[Link](https://arxiv.org/abs/2410.18351)Cited by:[§6](https://arxiv.org/html/2609.00355#S6.p1.1)\.
- Ankneret al\.\(2024\)Z\. Ankner, R\. Parthasarathy, A\. Nrusimha, C\. Rinard, J\. Ragan\-Kelley, and W\. BrandonHydra: sequentially\-dependent draft heads for Medusa decoding\.InConference on Language Modeling \(COLM\),Note:arXiv:2402\.05109Cited by:[§2](https://arxiv.org/html/2609.00355#S2.p1.1)\.
- AQ\-MedAI \(2025\)AQ\-MedAIQwen3\-VL\-8B\-Instruct\-EAGLE3: a production EAGLE\-3 draft head for Qwen3\-VL\-8B\.Note:https://huggingface\.co/AQ\-MedAI/Qwen3\-VL\-8B\-Instruct\-eagle3HuggingFace model checkpoint; run via SGLang native EAGLE\-3 treeCited by:[§1](https://arxiv.org/html/2609.00355#S1.p3.1),[§5\.1](https://arxiv.org/html/2609.00355#S5.SS1.SSS0.Px2.p1.1)\.
- Arriolaet al\.\(2025\)M\. Arriola, A\. Gokaslan, J\. T\. Chiu, Z\. Yang, Z\. Qi, J\. Han, S\. S\. Sahoo, and V\. KuleshovBlock Diffusion: Interpolating Between Autoregressive and Diffusion Language Models\.External Links:2503\.09573,[Link](https://arxiv.org/abs/2503.09573v3)Cited by:[§1](https://arxiv.org/html/2609.00355#S1.p5.1),[§2](https://arxiv.org/html/2609.00355#S2.p3.1)\.
- Baiet al\.\(2025\)S\. Bai, K\. Chen, X\. Liu, J\. Wang, W\. Ge, S\. Song, K\. Dang, P\. Wang, S\. Wang, J\. Tang,et al\.Qwen2\.5\-VL Technical Report\.arXiv preprint arXiv:2502\.13923\.Cited by:[§1](https://arxiv.org/html/2609.00355#S1.p1.1)\.
- Caiet al\.\(2024\)T\. Cai, Y\. Li, Z\. Geng, H\. Peng, J\. D\. Lee, D\. Chen, and T\. DaoMedusa: simple LLM inference acceleration framework with multiple decoding heads\.InInternational Conference on Machine Learning,Cited by:[1st item](https://arxiv.org/html/2609.00355#A4.I1.i1.p1.1),[§1](https://arxiv.org/html/2609.00355#S1.p2.1),[§2](https://arxiv.org/html/2609.00355#S2.p1.1)\.
- Chenet al\.\(2023\)C\. Chen, S\. Borgeaud, G\. Irving, J\. Lespiau, L\. Sifre, and J\. JumperAccelerating large language model decoding with speculative sampling\.arXiv preprint arXiv:2302\.01318\.Cited by:[§1](https://arxiv.org/html/2609.00355#S1.p2.1),[§2](https://arxiv.org/html/2609.00355#S2.p1.1),[§3\.1](https://arxiv.org/html/2609.00355#S3.SS1.p1.1)\.
- Chenet al\.\(2024a\)G\. H\. Chen, S\. Chen, R\. Zhang, J\. Chen, X\. Wu, Z\. Zhang, Z\. Chen, J\. Li, X\. Wan, and B\. WangALLaVA: harnessing GPT4V\-synthesized data for lite vision\-language models\.External Links:2402\.11684,[Link](https://arxiv.org/abs/2402.11684)Cited by:[Appendix D](https://arxiv.org/html/2609.00355#A4.SS0.SSS0.Px4.p1.1)\.
- Chenet al\.\(2026\)J\. Chen, Y\. Liang, and Z\. LiuDFlash: Block Diffusion for Flash Speculative Decoding\.External Links:2602\.06036,[Link](https://arxiv.org/abs/2602.06036v2)Cited by:[§B\.5](https://arxiv.org/html/2609.00355#A2.SS5.p1.1),[§1](https://arxiv.org/html/2609.00355#S1.p5.1),[§2](https://arxiv.org/html/2609.00355#S2.p3.1),[§3\.2](https://arxiv.org/html/2609.00355#S3.SS2.p1.1),[§4\.1](https://arxiv.org/html/2609.00355#S4.SS1.p1.1)\.
- Chenet al\.\(2024b\)Z\. Chen, A\. May, R\. Svirschevski, Y\. Huang, M\. Ryabinin, Z\. Jia, and B\. ChenSequoia: scalable, robust, and hardware\-aware speculative decoding\.InAdvances in Neural Information Processing Systems,Note:arXiv:2402\.12374Cited by:[§1](https://arxiv.org/html/2609.00355#S1.p2.1),[§2](https://arxiv.org/html/2609.00355#S2.p1.1)\.
- Chenget al\.\(2025\)Z\. Cheng, G\. Yang, J\. Li, Z\. Deng, M\. Guo, and S\. HuDEER: draft with diffusion, verify with autoregressive models\.External Links:2512\.15176,[Link](https://arxiv.org/abs/2512.15176)Cited by:[§2](https://arxiv.org/html/2609.00355#S2.p3.1)\.
- Christopheret al\.\(2025\)J\. K\. Christopher, B\. R\. Bartoldson, T\. Ben\-Nun, M\. Cardei, B\. Kailkhura, and F\. FiorettoSpeculative diffusion decoding: accelerating language generation through diffusion\.InProceedings of the North American Chapter of the Association for Computational Linguistics \(NAACL\),Note:arXiv:2408\.05636Cited by:[§2](https://arxiv.org/html/2609.00355#S2.p3.1)\.
- Fuet al\.\(2024\)Y\. Fu, P\. Bailis, I\. Stoica, and H\. ZhangBreak the sequential dependency of LLM inference using lookahead decoding\.InProceedings of the International Conference on Machine Learning \(ICML\),Note:arXiv:2402\.02057Cited by:[§2](https://arxiv.org/html/2609.00355#S2.p3.1)\.
- Gagraniet al\.\(2024\)M\. Gagrani, R\. Goel, W\. Jeon, J\. Park, M\. Lee, and C\. LottOn speculative decoding for multimodal large language models\.arXiv preprint arXiv:2404\.08856\.Note:ELVM @ CVPR 2024Cited by:[Appendix C](https://arxiv.org/html/2609.00355#A3.p1.1),[§1](https://arxiv.org/html/2609.00355#S1.p3.1),[§2](https://arxiv.org/html/2609.00355#S2.p2.1)\.
- Ganesanet al\.\(2025\)M\. Ganesan, S\. Segal, A\. Aggarwal, N\. Sinnadurai, S\. Lie, and V\. ThangarasaMASSV: multimodal adaptation and self\-data distillation for speculative decoding of vision\-language models\.Note:Findings of EMNLP 2025External Links:2505\.10526,[Link](https://arxiv.org/abs/2505.10526)Cited by:[§2](https://arxiv.org/html/2609.00355#S2.p2.1)\.
- Huet al\.\(2025\)Y\. Hu, T\. Xia, Z\. Liu, R\. Raman, X\. Liu, B\. Bao, E\. Sather, V\. Thangarasa, and S\. Q\. ZhangDREAM: drafting with refined target features and entropy\-adaptive cross\-attention fusion for multimodal speculative decoding\.InAdvances in Neural Information Processing Systems,Note:arXiv:2505\.19201Cited by:[§2](https://arxiv.org/html/2609.00355#S2.p2.1)\.
- Huanget al\.\(2025\)H\. Huang, F\. Yang, Z\. Liu, X\. Yin, D\. Li, P\. Ren, and E\. BarsoumSpecVLM: fast speculative decoding in vision\-language models\.External Links:2509\.11815,[Link](https://arxiv.org/abs/2509.11815)Cited by:[§1](https://arxiv.org/html/2609.00355#S1.p3.1),[§2](https://arxiv.org/html/2609.00355#S2.p2.1)\.
- Huanget al\.\(2024\)K\. Huang, X\. Guo, and M\. WangSpecDec\+\+: boosting speculative decoding via adaptive candidate lengths\.External Links:2405\.19715,[Link](https://arxiv.org/abs/2405.19715)Cited by:[§2](https://arxiv.org/html/2609.00355#S2.p1.1)\.
- Kanget al\.\(2025\)J\. Kang, H\. Shu, W\. Li, Y\. Zhai, and X\. ChenViSpec: accelerating vision\-language models with vision\-aware speculative decoding\.Note:NeurIPS 2025External Links:2509\.15235,[Link](https://arxiv.org/abs/2509.15235)Cited by:[§1](https://arxiv.org/html/2609.00355#S1.p3.1),[§2](https://arxiv.org/html/2609.00355#S2.p2.1),[§3\.1](https://arxiv.org/html/2609.00355#S3.SS1.p2.2),[§5\.1](https://arxiv.org/html/2609.00355#S5.SS1.SSS0.Px2.p1.1)\.
- Leviathanet al\.\(2023\)Y\. Leviathan, M\. Kalman, and Y\. MatiasFast inference from transformers via speculative decoding\.InInternational Conference on Machine Learning,Cited by:[§1](https://arxiv.org/html/2609.00355#S1.p2.1),[§2](https://arxiv.org/html/2609.00355#S2.p1.1),[§3\.1](https://arxiv.org/html/2609.00355#S3.SS1.p1.1)\.
- Liet al\.\(2025a\)G\. Li, Z\. Fu, M\. Fang, Q\. Zhao, M\. Tang, C\. Yuan, and J\. WangDiffuSpec: Unlocking Diffusion Language Models for Speculative Decoding\.External Links:2510\.02358,[Link](https://arxiv.org/abs/2510.02358v1)Cited by:[§2](https://arxiv.org/html/2609.00355#S2.p3.1)\.
- Liet al\.\(2024a\)Y\. Li, F\. Wei, C\. Zhang, and H\. ZhangEAGLE\-2: faster inference of language models with dynamic draft trees\.InEmpirical Methods in Natural Language Processing,Cited by:[§1](https://arxiv.org/html/2609.00355#S1.p2.1),[§2](https://arxiv.org/html/2609.00355#S2.p1.1),[§5\.1](https://arxiv.org/html/2609.00355#S5.SS1.SSS0.Px2.p1.1)\.
- Liet al\.\(2024b\)Y\. Li, F\. Wei, C\. Zhang, and H\. ZhangEAGLE: speculative sampling requires rethinking feature uncertainty\.InInternational Conference on Machine Learning,Cited by:[§1](https://arxiv.org/html/2609.00355#S1.p2.1),[§2](https://arxiv.org/html/2609.00355#S2.p1.1),[§5\.1](https://arxiv.org/html/2609.00355#S5.SS1.SSS0.Px2.p1.1)\.
- Liet al\.\(2025b\)Y\. Li, F\. Wei, C\. Zhang, and H\. ZhangEAGLE\-3: scaling up inference acceleration of large language models via training\-time test\.InAdvances in Neural Information Processing Systems,Note:arXiv:2503\.01840Cited by:[§1](https://arxiv.org/html/2609.00355#S1.p2.1),[§2](https://arxiv.org/html/2609.00355#S2.p1.1),[§4\.1](https://arxiv.org/html/2609.00355#S4.SS1.p1.1),[§5\.1](https://arxiv.org/html/2609.00355#S5.SS1.SSS0.Px2.p1.1)\.
- Linet al\.\(2014\)T\. Lin, M\. Maire, S\. Belongie, J\. Hays, P\. Perona, D\. Ramanan, P\. Dollár, and C\. L\. ZitnickMicrosoft COCO: common objects in context\.InECCV,Cited by:[§5\.1](https://arxiv.org/html/2609.00355#S5.SS1.SSS0.Px1.p1.1)\.
- Luet al\.\(2022\)P\. Lu, S\. Mishra, T\. Xia, L\. Qiu, K\. Chang, S\. Zhu, O\. Tafjord, P\. Clark, and A\. KalyanLearn to explain: multimodal reasoning via thought chains for science question answering\.InAdvances in Neural Information Processing Systems,Cited by:[§5\.4](https://arxiv.org/html/2609.00355#S5.SS4.p3.1)\.
- Mahmoud \(2026\)S\. MahmoudAcceptance dynamics across cognitive domains in speculative decoding\.External Links:2604\.14682,[Link](https://arxiv.org/abs/2604.14682)Cited by:[§6](https://arxiv.org/html/2609.00355#S6.p1.1)\.
- Masryet al\.\(2022\)A\. Masry, D\. X\. Long, J\. Q\. Tan, S\. Joty, and E\. HoqueChartQA: a benchmark for question answering about charts with visual and logical reasoning\.InFindings of ACL,Cited by:[§5\.1](https://arxiv.org/html/2609.00355#S5.SS1.SSS0.Px1.p1.1)\.
- Mathewet al\.\(2022\)M\. Mathew, V\. Bagal, R\. Tito, D\. Karatzas, E\. Valveny, and C\. JawaharInfographicVQA\.InWACV,Cited by:[§5\.1](https://arxiv.org/html/2609.00355#S5.SS1.SSS0.Px1.p1.1)\.
- Mathewet al\.\(2021\)M\. Mathew, D\. Karatzas, and C\. JawaharDocVQA: a dataset for VQA on document images\.InWACV,Cited by:[§5\.1](https://arxiv.org/html/2609.00355#S5.SS1.SSS0.Px1.p1.1)\.
- Miaoet al\.\(2024\)X\. Miao, G\. Oliaro, Z\. Zhang, X\. Cheng, Z\. Wang, Z\. Zhang, R\. Y\. Y\. Wong, A\. Zhu, L\. Yang, X\. Shi, C\. Shi, Z\. Chen, D\. Arfeen, R\. Abhyankar, and Z\. JiaSpecInfer: accelerating large language model serving with tree\-based speculative inference and verification\.InASPLOS,Cited by:[§1](https://arxiv.org/html/2609.00355#S1.p2.1),[§2](https://arxiv.org/html/2609.00355#S2.p1.1),[§4\.2](https://arxiv.org/html/2609.00355#S4.SS2.p1.1)\.
- Moneaet al\.\(2023\)G\. Monea, A\. Joulin, and E\. GravePaSS: parallel speculative sampling\.arXiv preprint arXiv:2311\.13581\.Note:NeurIPS 2023 Workshop on Efficient Natural Language and Speech ProcessingCited by:[§2](https://arxiv.org/html/2609.00355#S2.p3.1)\.
- Nieet al\.\(2025\)S\. Nie, F\. Zhu, Z\. You, X\. Zhang, J\. Ou, J\. Hu, J\. Zhou, Y\. Lin, J\. Wen, and C\. LiLarge Language Diffusion Models\.External Links:2502\.09992,[Link](https://arxiv.org/abs/2502.09992v3)Cited by:[§2](https://arxiv.org/html/2609.00355#S2.p3.1)\.
- Qwen Team \(2025\)Qwen TeamQwen3\-VL Technical Report\.Note:https://huggingface\.co/Qwen/Qwen3\-VL\-8B\-InstructQwen3\-VL\-8B\-InstructCited by:[§1](https://arxiv.org/html/2609.00355#S1.p1.1),[§5\.1](https://arxiv.org/html/2609.00355#S5.SS1.SSS0.Px1.p1.1)\.
- Ringel and Romano \(2026\)L\. Ringel and Y\. RomanoAccelerating Speculative Decoding with Block Diffusion Draft Trees\.External Links:2604\.12989,[Link](https://arxiv.org/abs/2604.12989v1)Cited by:[§B\.5](https://arxiv.org/html/2609.00355#A2.SS5.p1.1),[§1](https://arxiv.org/html/2609.00355#S1.p5.1),[§2](https://arxiv.org/html/2609.00355#S2.p3.1),[§3\.2](https://arxiv.org/html/2609.00355#S3.SS2.p1.1),[§4\.2](https://arxiv.org/html/2609.00355#S4.SS2.p2.1)\.
- SGLang Team \(2026\)SGLang TeamSpecForge: a flexible and efficient open\-source training framework for speculative decoding\.arXiv preprint arXiv:2603\.18567\.Cited by:[§5\.3](https://arxiv.org/html/2609.00355#S5.SS3.p1.1)\.
- Shenet al\.\(2026\)H\. Shen, X\. Wang, P\. Zhang, Y\. Hsieh, Q\. Han, Z\. Wan, Z\. Zhang, J\. Zhang, J\. Xiong, Z\. Liu, Y\. Zhang, H\. Cao, C\. Zhao, and M\. ZhangMMSpec: benchmarking speculative decoding for vision\-language models\.External Links:2603\.14989,[Link](https://arxiv.org/abs/2603.14989)Cited by:[§2](https://arxiv.org/html/2609.00355#S2.p2.1)\.
- Singhet al\.\(2019\)A\. Singh, V\. Natarajan, M\. Shah, Y\. Jiang, X\. Chen, D\. Batra, D\. Parikh, and M\. RohrbachTowards VQA models that can read\.InCVPR,Cited by:[§5\.1](https://arxiv.org/html/2609.00355#S5.SS1.SSS0.Px1.p1.1)\.
- Tonget al\.\(2026\)Y\. Tong, T\. Zhang, Y\. Wan, K\. Lin, J\. Yuan, and C\. HuSAGE: accelerating vision\-language models via entropy\-guided adaptive speculative decoding\.External Links:2602\.00523,[Link](https://arxiv.org/abs/2602.00523)Cited by:[§6](https://arxiv.org/html/2609.00355#S6.p1.1)\.
- Wanget al\.\(2025a\)J\. Wang, Y\. Su, J\. Li, Q\. Xia, Z\. Ye, X\. Duan, Z\. Wang, and M\. ZhangOPT\-Tree: speculative decoding with adaptive draft tree structure\.Transactions of the Association for Computational Linguistics \(TACL\)\.Note:arXiv:2406\.17276Cited by:[§2](https://arxiv.org/html/2609.00355#S2.p1.1)\.
- Wanget al\.\(2025b\)Z\. Wang, R\. Li, H\. Du, J\. T\. Zhou, Y\. Zhang, and X\. YangSpecFLASH: a latent\-guided semi\-autoregressive speculative decoding framework for efficient multimodal generation\.External Links:2505\.12728,[Link](https://arxiv.org/abs/2505.12728)Cited by:[§2](https://arxiv.org/html/2609.00355#S2.p2.1)\.
- Wuet al\.\(2026\)C\. Wu, S\. Lan, Y\. Fu, S\. Gao, J\. Wang, J\. Yu, J\. M\. Alvarez, P\. Molchanov, P\. Luo, S\. Han, L\. Zhu, and E\. XieFast\-dVLM: efficient block\-diffusion vlm via direct conversion from autoregressive vlm\.arXiv preprint arXiv:2604\.06832\.External Links:[Link](https://arxiv.org/abs/2604.06832)Cited by:[§2](https://arxiv.org/html/2609.00355#S2.p3.1)\.
- Xiaet al\.\(2024\)H\. Xia, Z\. Yang, Q\. Dong, P\. Wang, Y\. Li, T\. Ge, T\. Liu, W\. Li, and Z\. SuiUnlocking efficiency in large language model inference: a comprehensive survey of speculative decoding\.InFindings of the Association for Computational Linguistics \(ACL Findings\),Note:Spec\-BenchCited by:[§2](https://arxiv.org/html/2609.00355#S2.p1.1)\.
- Xieet al\.\(2025\)Z\. Xie, P\. Wang, S\. Qiu, and J\. ChengHiViS: hiding visual tokens from the drafter for speculative decoding in vision\-language models\.External Links:2509\.23928,[Link](https://arxiv.org/abs/2509.23928)Cited by:[§1](https://arxiv.org/html/2609.00355#S1.p3.1),[§2](https://arxiv.org/html/2609.00355#S2.p2.1)\.
- Yinet al\.\(2024\)M\. Yin, M\. Chen, K\. Huang, and M\. WangA theoretical perspective for speculative decoding algorithm\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Note:arXiv:2411\.00841Cited by:[§2](https://arxiv.org/html/2609.00355#S2.p1.1)\.
- Zhanget al\.\(2026a\)K\. Zhang, J\. Wang, S\. Gao, C\. Wu, Y\. Cao, S\. Han, B\. Ivanovic, L\. Liu, M\. Pavone, S\. Han, D\. Zhou, and E\. XieFast\-dDrive: efficient block\-diffusion VLM for autonomous driving\.External Links:2605\.23163,[Link](https://arxiv.org/abs/2605.23163)Cited by:[§2](https://arxiv.org/html/2609.00355#S2.p3.1)\.
- Zhanget al\.\(2026b\)S\. Zhang, H\. Qiu, H\. He, and Y\. DaiCost\-Aware Diffusion Draft Trees for Speculative Decoding\.External Links:2606\.01813,[Link](https://arxiv.org/abs/2606.01813)Cited by:[§B\.5](https://arxiv.org/html/2609.00355#A2.SS5.p1.1),[§1](https://arxiv.org/html/2609.00355#S1.p5.1),[§2](https://arxiv.org/html/2609.00355#S2.p3.1),[§4\.2](https://arxiv.org/html/2609.00355#S4.SS2.p2.1)\.

Technical Appendix

## Appendix ALimitations

All claims are batch one and decode\-only, and we make no larger\-batch claim\. The bitwise guarantee is a greedy\-regime statement: the sampled\-decoding results of Section[5\.4](https://arxiv.org/html/2609.00355#S5.SS4)preserve the target’s distribution by construction, not byte\-identity, and are reported on one run at each temperature\. Sample sizes are small, and they are stated at each measurement in Appendices[D](https://arxiv.org/html/2609.00355#A4)and[E](https://arxiv.org/html/2609.00355#A5)\. The bitwise guarantee is stated in fp32\. Under production bf16 it holds as byte\-identity through a tie\-aware gate\. Two of the three cross\-target fits use two\-model proxies rather than trained block adapters\. The grounded ordering itself is architecture\-scoped, and it inverts under cross\-attention fusion\.

## Appendix BTheory: Statements and Proofs

### B\.1Losslessness

###### Theorem[1](https://arxiv.org/html/2609.00355#Thmtheorem1)\(full\)\.

For the decoder of Definition[1](https://arxiv.org/html/2609.00355#Thmdefinition1), assume \(float\-exactness\) that at every visited prefixx∘b∘y1:kx\\circ b\\circ y\_\{1:k\}the targetarg​max\\argmaxis unique and the top logit gapγk:=ℓ\(1\)−ℓ\(2\)\>0\\gamma\_\{k\}:=\\ell\_\{\(1\)\}\-\\ell\_\{\(2\)\}\>0exceeds the packed\-versus\-unpacked logit perturbation\. Then for any candidate tree the committed token sequence equals the target’s autoregressive greedy sequence exactly, and losslessness can fail only at a marginγk\\gamma\_\{k\}below that perturbation, a tie\.

###### Proof\.

*Single round\.*By Definition[1](https://arxiv.org/html/2609.00355#Thmdefinition1)the verifier row at any prefixx∘b∘y1:kx\\circ b\\circ y\_\{1:k\}reproducesp\(⋅∣x∘b∘y1:k\)p\(\\cdot\\mid x\\circ b\\circ y\_\{1:k\}\)exactly, since ancestor\-only masking and token packing change the attention layout, not the row\-level conditioning\. LetYk\+1Y\_\{k\+1\}be the target greedy token there\. The walk then takes one of two branches:

y1:k∘Yk\+1∈S\\displaystyle y\_\{1:k\}\\circ Y\_\{k\+1\}\\in S:descend and commit​Yk\+1;\\displaystyle:\\ \\text\{descend and commit \}Y\_\{k\+1\};otherwise\\displaystyle\\text\{otherwise\}:stop and emit​Yk\+1​as the next root\.\\displaystyle:\\ \\text\{stop and emit \}Y\_\{k\+1\}\\text\{ as the next root\.\}Either way the token produced at that position isYk\+1Y\_\{k\+1\}, and the tree affects only how many tokens commit for the block\.

*Induction over rounds\.*The emitted stream concatenates accepted paths and corrections, and each round’s correction is the next round’s root\. Every produced token is the target greedy token at its own prefix, so induction on output position gives the committed sequence

Y1Y2⋯bitwise, independent of all budget choices\.Y\_\{1\}Y\_\{2\}\\cdots\\qquad\\text\{bitwise, independent of all budget choices\.\}
*Tie sensitivity\.*The only breakable step is the equality of the packed and autoregressivearg​max\\argmax\. It fails only at

γk<\(packed\-versus\-unpacked logit perturbation\),\\gamma\_\{k\}<\\text\{\(packed\-versus\-unpacked logit perturbation\)\},a near\-tie\. This is the operational bf16 tie\-aware gate of Section[4\.2](https://arxiv.org/html/2609.00355#S4.SS2): gating rounds whose observed margins fall under the threshold keeps the perturbation bound in force, and the fp32 audit shows zero divergences\. ∎

### B\.2Derivation of the Law

###### Assumption 2\(Offset\-wise survival with marginal hazards\)\.

Conditioned on the round \(equivalently onHH\), let the drafter top\-11token match the target greedy token at offsetjjwith probabilitypm,jp\_\{\\mathrm\{m\},j\}, and letaabe the leading run of matches before the first miss, capped atLL\. We assume the survival factorizes into the marginal hazards,

ℙ⁡\[a≥ℓ∣H\]=∏j≤ℓpm,j,𝔼⁡\[a∣H\]=∑ℓ=1L∏j≤ℓpm,j\.\\mathbb\{P\}\[a\\geq\\ell\\mid H\]=\\prod\_\{j\\leq\\ell\}p\_\{\\mathrm\{m\},j\},\\qquad\\mathbb\{E\}\[a\\mid H\]=\\sum\_\{\\ell=1\}^\{L\}\\prod\_\{j\\leq\\ell\}p\_\{\\mathrm\{m\},j\}\.\(3\)This is an assumption, not a consequence of the chain rule: it replaces the conditional hazardsℙ\[matchj∣match<j,H\]\\mathbb\{P\}\[\\mathrm\{match\}\_\{j\}\\mid\\mathrm\{match\}\_\{<j\},H\]by the marginalspm,jp\_\{\\mathrm\{m\},j\}\.

###### Lemma 1\(Slowly varying hazard gives a truncated\-geometric mean\)\.

Under Assumption[2](https://arxiv.org/html/2609.00355#Thmassumption2)and the further approximationpm,j≈pmp\_\{\\mathrm\{m\},j\}\\approx p\_\{\\mathrm\{m\}\}for allj≤Lj\\leq L, Equation \([3](https://arxiv.org/html/2609.00355#A2.E3)\) reduces to𝔼⁡\[a∣H\]≈∑ℓ=1Lpmℓ=pm​\(1−pmL\)/\(1−pm\)→pm/\(1−pm\)\\mathbb\{E\}\[a\\mid H\]\\approx\\sum\_\{\\ell=1\}^\{L\}p\_\{\\mathrm\{m\}\}^\{\\ell\}=p\_\{\\mathrm\{m\}\}\(1\-p\_\{\\mathrm\{m\}\}^\{L\}\)/\(1\-p\_\{\\mathrm\{m\}\}\)\\to p\_\{\\mathrm\{m\}\}/\(1\-p\_\{\\mathrm\{m\}\}\)asL→∞L\\to\\infty, the untruncated form being an upper bound for all finiteLLwith relative truncation errorpmLp\_\{\\mathrm\{m\}\}^\{L\}\.

###### Proof\.

Settingpm,j≡pmp\_\{\\mathrm\{m\},j\}\\equiv p\_\{\\mathrm\{m\}\}in Equation \([3](https://arxiv.org/html/2609.00355#A2.E3)\) gives the finite geometric series\.

Its truncated tail is

∑ℓ\>Lpmℓ=pmL⋅pm1−pm,\\sum\_\{\\ell\>L\}p\_\{\\mathrm\{m\}\}^\{\\ell\}=p\_\{\\mathrm\{m\}\}^\{L\}\\cdot\\frac\{p\_\{\\mathrm\{m\}\}\}\{1\-p\_\{\\mathrm\{m\}\}\},so the relative truncation error ispmLp\_\{\\mathrm\{m\}\}^\{L\}\.

AtL=15L\{=\}15:

pmL=0\.002,0\.035,0\.206atpm=0\.65,0\.80,0\.90,p\_\{\\mathrm\{m\}\}^\{L\}=0\.002,\\ 0\.035,\\ 0\.206\\quad\\text\{at\}\\quad p\_\{\\mathrm\{m\}\}=0\.65,\\ 0\.80,\\ 0\.90,so the closed form is tight except near the entropy floor\.

Two errors are folded in: the truncation tail, and the slowly\-varying step itself\. The step overstates𝔼⁡\[a∣H\]\\mathbb\{E\}\[a\\mid H\]whenpm,jp\_\{\\mathrm\{m\},j\}decays injj\. It understates the extreme\-low\-entropy mean when onepmp\_\{\\mathrm\{m\}\}is fit across a range containing a sharp low\-HHspike: the fit’s supremum is3\.343\.34atb0=1\.226b\_\{0\}\{=\}1\.226,L=15L\{=\}15, below the measured DocVQA bottom\-decile mean of5\.235\.23\(see Theorem[2](https://arxiv.org/html/2609.00355#Thmtheorem2)\)\.

Substitutingpm​\(H\)=σ⁡\(b0−b1​H\)p\_\{\\mathrm\{m\}\}\(H\)=\\sigma\(b\_\{0\}\-b\_\{1\}H\)gives Equation \([2](https://arxiv.org/html/2609.00355#S6.E2)\), which we use as an operational conditional\-mean characterization, not an exact survival identity\. ∎

###### Proposition 1\(Entropy controls the match, and the affine logit is a surrogate\)\.

Letpmp\_\{\\mathrm\{m\}\}be the probability that the drafter’s root top\-11equals the target greedy token, andHHthe target’s root entropy\. \(i\) Rigorously, the target’s own top\-11errore⋆=1−maxw⁡p⁡\(w∣x,b\)e^\{\\star\}=1\-\\max\_\{w\}p\(w\\mid x,b\)obeys Fano’s inequality

H≤ℍb​\(e⋆\)\+e⋆​log⁡\(\|Σ\|−1\),H\\leq\\mathbb\{H\}\_\{b\}\(e^\{\\star\}\)\+e^\{\\star\}\\log\(\|\\Sigma\|\-1\),so largerHHforces largere⋆e^\{\\star\}andH→0H\\to 0forcesmaxw⁡p→1\\max\_\{w\}p\\to 1\. \(ii\) Modeling the drafter’s logit error as logistic with scalessgivespm=σ⁡\(Δ/s\)p\_\{\\mathrm\{m\}\}=\\sigma\(\\Delta/s\)for the target top\-22marginΔ⁡\(H\)\\Delta\(H\), decreasing inHHby \(i\)\. \(iii\) A first\-order expansion ofΔ\\DeltainHHgiveslogit⁡pm≈b0−b1​H\\operatorname\{logit\}p\_\{\\mathrm\{m\}\}\\approx b\_\{0\}\-b\_\{1\}Hwithb1=−Δ′\(H0\)/s\>0b\_\{1\}=\-\\Delta^\{\\prime\}\(H\_\{0\}\)/s\>0\. Part \(i\) predicts sign and monotonicity\. Parts \(ii\) and \(iii\) are modeling steps that select the functional form\.

###### Proof\.

1. \(i\)View the target top\-11token as an estimator of a drawW∼pW\\sim p\. Fano’s inequality boundsHHby an increasing function of the errore⋆e^\{\\star\}on\[0,1−1/\|Σ\|\]\[0,1\-1/\|\\Sigma\|\], soe⋆e^\{\\star\}grows withHHand vanishes asH→0H\\to 0\.
2. \(ii\)The drafter preserves the target ordering iff the perturbed margin stays positive\. A difference of logistics is logistic, so logit⁡pm=Δ/s,\\operatorname\{logit\}p\_\{\\mathrm\{m\}\}=\\Delta/s,and on a fixed\-entropy shellΔ\\Deltais decreasing inHHby \(i\)\.
3. \(iii\)This is the first\-order expansion\. Its adequacy over the empirical window is what the curveR2R^\{2\}measures\.

A fuller derivation under a temperature family \(target logits a fixed shape scaled by inverse temperature\) showslogit⁡p\(1\)tgt​\(H\)\\operatorname\{logit\}p^\{\\mathrm\{tgt\}\}\_\{\(1\)\}\(H\)is affine toR2=0\.84R^\{2\}=0\.84to0\.900\.90over the resolvable windowH∈\[0\.05,1\.2\]H\\in\[0\.05,1\.2\]for Zipf, geometric\-gap, and near\-binary logit shapes \(self\-confidence slopesκ≈3\.2\\kappa\\approx 3\.2to3\.93\.9\)\.

It reduces the match slope to a single drafter\-noise scaless, distinct from the acceptance lengthτ\\tauof Equation \([1](https://arxiv.org/html/2609.00355#S3.E1)\), viab1=κ/sb\_\{1\}=\\kappa/s\. Matching the measured offset\-11slopes gives

s=3\.4/7\.7/3\.8on captioning/TextVQA/DocVQA\.s=3\.4/7\.7/3\.8\\quad\\text\{on captioning/TextVQA/DocVQA\.\}The affine form is exactly a local linearization: the binary closed form is provably curved globally, so the law is stated over the empirical window only\. ∎

###### Corollary \(Grounded is more draftable, full\)\.

Under Equation \([2](https://arxiv.org/html/2609.00355#S6.E2)\),

∂Hpm\\displaystyle\\partial\_\{H\}p\_\{\\mathrm\{m\}\}=−b1​pm​\(1−pm\),\\displaystyle=\-b\_\{1\}p\_\{\\mathrm\{m\}\}\(1\-p\_\{\\mathrm\{m\}\}\),∂H𝔼⁡\[a∣H\]\\displaystyle\\partial\_\{H\}\\mathbb\{E\}\[a\\mid H\]=∂Hpm/\(1−pm\)2=−b1​𝔼​\[a∣H\]<0,\\displaystyle=\\partial\_\{H\}p\_\{\\mathrm\{m\}\}/\(1\-p\_\{\\mathrm\{m\}\}\)^\{2\}=\-b\_\{1\}\\mathbb\{E\}\[a\\mid H\]<0,so𝔼⁡\[a∣H\]\\mathbb\{E\}\[a\\mid H\]is strictly decreasing inHH\. Hence any regime with systematically lower next\-token entropy has a longer expected accepted block\. Grounded and OCR generation copies determinate glyph and answer strings from the image, collapsing the target’s entropy, so grounded spans are more draftable than open text and task\-level mean accepted length is ordered by grounding\.

###### Proof\.

The derivative computation is displayed in the statement and uses onlyb1\>0b\_\{1\}\>0andpm∈\(0,1\)p\_\{\\mathrm\{m\}\}\\in\(0,1\)\. The ordering claim uses only the measured monotone decrease ofaainHH, not the fitted closed form\.

The measured instantiation orders mean accepted length by grounding:

captioning​1\.88<TextVQA​2\.12<DocVQA​3\.08\.\\text\{captioning \}1\.88\\ <\\ \\text\{TextVQA \}2\.12\\ <\\ \\text\{DocVQA \}3\.08\.The later round logs extend the same ordering,2\.672\.67on InfoVQA and, under its templated prompts,3\.713\.71on ChartQA\. The gapDocVQA−captioning=\+1\.20\\text\{DocVQA\}\-\\text\{captioning\}=\+1\.20has95%95\\%bootstrap CI\[1\.07,1\.33\]\[1\.07,1\.33\]and Mann\-Whitneyp=8\.7×10−68p=8\.7\\times 10^\{\-68\}, and the grounded task carries the higherb0b\_\{0\}and steeperb1b\_\{1\}\(Table[12](https://arxiv.org/html/2609.00355#A6.T12)\)\. ∎

### B\.3The Noise Ceiling

###### Proposition 2\(Entropy\-only noise ceiling\)\.

Among allσ⁡\(H\)\\sigma\(H\)\-measurable predictorsg⁡\(H\)g\(H\)ofaa, the largest attainable coefficient of determination isR¯2=Var⁡\(𝔼⁡\[a∣H\]\)/Var⁡\(a\)=1−𝔼⁡\[Var⁡\(a∣H\)\]/Var⁡\(a\)\\bar\{R\}^\{2\}=\\mathrm\{Var\}\(\\mathbb\{E\}\[a\\mid H\]\)/\\mathrm\{Var\}\(a\)=1\-\\mathbb\{E\}\[\\mathrm\{Var\}\(a\\mid H\)\]/\\mathrm\{Var\}\(a\), attained by the conditional mean\. For a truncated\-geometricaathe within\-round variance is of order𝔼​\[a∣H\]2\\mathbb\{E\}\[a\\mid H\]^\{2\}, soR¯2\\bar\{R\}^\{2\}is intrinsically small regardless of predictor quality\.

###### Proof\.

For anygg, conditional\-mean orthogonality gives

𝔼⁡\[\(a−g⁡\(H\)\)2\]=𝔼⁡\[Var⁡\(a∣H\)\]\+𝔼⁡\[\(𝔼⁡\[a∣H\]−g⁡\(H\)\)2\],\\mathbb\{E\}\[\(a\-g\(H\)\)^\{2\}\]=\\mathbb\{E\}\[\\mathrm\{Var\}\(a\\mid H\)\]\+\\mathbb\{E\}\[\(\\mathbb\{E\}\[a\\mid H\]\-g\(H\)\)^\{2\}\],minimized atg=𝔼⁡\[a∣H\]g=\\mathbb\{E\}\[a\\mid H\]\. The law of total variance then gives the stated ratio\.

For the untruncated geometric, withμ=𝔼⁡\[a∣H\]\\mu=\\mathbb\{E\}\[a\\mid H\],

Var⁡\(a∣H\)=μ⁡\(1\+μ\)≥μ2,\\mathrm\{Var\}\(a\\mid H\)=\\mu\(1\+\\mu\)\\geq\\mu^\{2\},so within\-round variance dominates unlessVar⁡\(μ⁡\(H\)\)\\mathrm\{Var\}\(\\mu\(H\)\)is comparably large\.

A nonparametric5050\-bin between/total variance estimate gives ceilings

0\.110/0\.144/0\.188on captioning/TextVQA/DocVQA\.0\.110/0\.144/0\.188\\quad\\text\{on captioning/TextVQA/DocVQA\.\}The law’s round\-levelR2=0\.097/0\.091/0\.074R^\{2\}=0\.097/0\.091/0\.074therefore realizes89%/63%/39%89\\%/63\\%/39\\%of what any entropy\-only predictor could, and the two later logs realize62%62\\%on InfoVQA and50%50\\%on ChartQA \(Table[12](https://arxiv.org/html/2609.00355#A6.T12)\), which is why the curveR2R^\{2\}is the correct score\. ∎

### B\.4The Grounded Tail

###### Assumption 3\(Two\-state determinate or uncertain survival\)\.

Conditioned on the round, offset\-wise matches follow a two\-state hidden Markov chain oversj∈\{D,U\}s\_\{j\}\\in\\\{D,U\\\}: entry distribution\(π,1−π\)\(\\pi,1\-\\pi\), match probabilityρs\\rho\_\{s\}in statess, transition matrixTTupon a match, run ending at the first miss, capped atLL\. StateDDmodels verbatim copying \(ρD→1\\rho\_\{D\}\\to 1, stickytD​D\>0t\_\{DD\}\>0\)\.UUis the ordinary moderate\-hazard regime\. The parameters governingH→0H\\to 0are estimated in regime, from near\-zero\-entropy rounds only, not extrapolated from the affine fit\. The geometric of Lemma[1](https://arxiv.org/html/2609.00355#Thmlemma1)is the degenerate caseρD=ρU\\rho\_\{D\}=\\rho\_\{U\}\.

###### Theorem[2](https://arxiv.org/html/2609.00355#Thmtheorem2)\(full\)\.

LetM⁡\(p\)=∑ℓ=1LpℓM\(p\)=\\sum\_\{\\ell=1\}^\{L\}p^\{\\ell\}\. \(A\) Under the single\-hazard law \([2](https://arxiv.org/html/2609.00355#S6.E2)\),supH≥0𝔼⁡\[a∣H\]=M⁡\(σ⁡\(b0\)\)≤eb0\\sup\_\{H\\geq 0\}\\mathbb\{E\}\[a\\mid H\]=M\(\\sigma\(b\_\{0\}\)\)\\leq e^\{b\_\{0\}\}: a rigorous, fit\-free cap\. \(B\) Under Assumption[3](https://arxiv.org/html/2609.00355#Thmassumption3), with𝐯=\(π,1−π\)\\mathbf\{v\}=\(\\pi,1\-\\pi\),D𝝆=diag⁡\(ρD,ρU\)D\_\{\\boldsymbol\{\\rho\}\}=\\operatorname\{diag\}\(\\rho\_\{D\},\\rho\_\{U\}\), andK=T​D𝝆K=TD\_\{\\boldsymbol\{\\rho\}\},

ℙ⁡\[a≥ℓ∣H\]\\displaystyle\\mathbb\{P\}\[a\\geq\\ell\\mid H\]=𝐯​D𝝆​Kℓ−1​𝟏,\\displaystyle=\\mathbf\{v\}D\_\{\\boldsymbol\{\\rho\}\}K^\{\\ell\-1\}\\mathbf\{1\},\(4\)𝔼⁡\[a∣H\]\\displaystyle\\mathbb\{E\}\[a\\mid H\]=𝐯​D𝝆​\(∑m=0L−1Km\)​𝟏,\\displaystyle=\\mathbf\{v\}D\_\{\\boldsymbol\{\\rho\}\}\\Big\(\\textstyle\\sum\_\{m=0\}^\{L\-1\}K^\{m\}\\Big\)\\mathbf\{1\},whose effective hazardhℓh\_\{\\ell\}is non\-constant, and takingρD,tD​D→1\\rho\_\{D\},t\_\{DD\}\\to 1gives𝔼​\[a∣H\]H→0→π​L\+\(1−π\)​M​\(ρU\)\\mathbb\{E\}\[a\\mid H\]^\{H\\to 0\}\\to\\pi L\+\(1\-\\pi\)M\(\\rho\_\{U\}\), which exceedseb0e^\{b\_\{0\}\}exactly whenπ\>π⋆:=\(eb0−M⁡\(ρU\)\)/\(L−M⁡\(ρU\)\)\\pi\>\\pi^\{\\star\}:=\(e^\{b\_\{0\}\}\-M\(\\rho\_\{U\}\)\)/\(L\-M\(\\rho\_\{U\}\)\)\. This threshold lies in\(0,1\)\(0,1\)becauseM⁡\(ρU\)<eb0<LM\(\\rho\_\{U\}\)<e^\{b\_\{0\}\}<L\(hereL=15L\{=\}15,eb0=3\.41e^\{b\_\{0\}\}\{=\}3\.41\)\. \(C\) With in\-regime maximum\-likelihood parameters, Equation \([4](https://arxiv.org/html/2609.00355#A2.E4)\) recovers the measured DocVQA floor mean5\.235\.23versus the geometric cap3\.413\.41\. The falsifiable validation is Proposition[3](https://arxiv.org/html/2609.00355#Thmproposition3)\.

###### Proof\.

1. \(A\)MMis strictly increasing on\(0,1\)\(0,1\), andpm​\(H\)p\_\{\\mathrm\{m\}\}\(H\)is maximized asH→0H\\to 0where it equalsσ⁡\(b0\)\\sigma\(b\_\{0\}\)\. ThenM⁡\(p\)≤p/\(1−p\)M\(p\)\\leq p/\(1\-p\)gives theeb0e^\{b\_\{0\}\}bound\.
2. \(B\)Starting from𝐯\\mathbf\{v\}, the joint event “match at offset11, now in statess” has mass\(𝐯​D𝝆\)s\(\\mathbf\{v\}D\_\{\\boldsymbol\{\\rho\}\}\)\_\{s\}, and each accepted offset multiplies byK=T​D𝝆K=TD\_\{\\boldsymbol\{\\rho\}\}, so ℙ⁡\[a≥ℓ∣H\]=𝐯​D𝝆​Kℓ−1​𝟏\\mathbb\{P\}\[a\\geq\\ell\\mid H\]=\\mathbf\{v\}D\_\{\\boldsymbol\{\\rho\}\}K^\{\\ell\-1\}\\mathbf\{1\}and the tail\-sum identity gives the mean\. Survival is a nonnegative combination of the two eigenmodes ofKK, so the hazardhℓh\_\{\\ell\}is monotone towardλmax​\(K\)\\lambda\_\{\\max\}\(K\)and non\-constant unlessρD=ρU\\rho\_\{D\}=\\rho\_\{U\}\. Its first valueπ​ρD\+\(1−π\)​ρU\\pi\\rho\_\{D\}\+\(1\-\\pi\)\\rho\_\{U\}can approach11and exceedσ⁡\(b0\)\\sigma\(b\_\{0\}\), which the constant\-hazard family cannot\. AsρD,tD​D→1\\rho\_\{D\},t\_\{DD\}\\to 1, ℙ\[a≥ℓ\]→πfor allℓ≤L,\\mathbb\{P\}\[a\\geq\\ell\]\\to\\pi\\quad\\text\{for all \}\\ell\\leq L,giving the stated cap violation\.
3. \(C\)Numerical instantiation, recovering the measured floor mean against the geometric cap\.∎

###### Proposition 3\(Prompt\-disjoint held\-out falsification\)\.

On the122122bottom\-entropy\-decile DocVQA rounds \(4747prompts\), split by prompt into disjoint sets: \(i\) on the full floor the geometric run\-length is rejected \(Pearsonχ2=105\.3\\chi^\{2\}=105\.3,df=7\\mathrm\{df\}=7,p=8\.7×10−20p=8\.7\\times 10^\{\-20\}\) while the likelihood ratio decisively favors the two\-state form \(2​Δ​ℓ=86\.52\\Delta\\ell=86\.5,df=4\\mathrm\{df\}=4,p=7\.2×10−18p=7\.2\\times 10^\{\-18\}\)\. The two\-state fit is itself also rejected in absolute terms \(χ2=29\.6\\chi^\{2\}=29\.6,df=3\\mathrm\{df\}=3\)\. \(ii\) Frozen on one prompt set and scored on the disjoint one, the two\-state model attains a smaller held\-outχ2\\chi^\{2\}than the geometric in66of77tests \(both split directions and five CV folds, mean held\-outχ2\\chi^\{2\}4\.44\.4against11\.911\.9\)\. \(iii\) Absolutely, the two\-state survives the held\-out shape test in one split direction \(χ2=6\.1\\chi^\{2\}=6\.1,p=0\.52p=0\.52\) and four of five folds, while the geometric survives in neither direction\. The load\-bearing claim is the relative held\-out dominance \(ii\)\. We do not claim the two\-state law is the exact data\-generating process\.

###### Proof\.

The proposition records the outcome of the stated estimation procedure\. Construction and statistics are as listed\.

The frozen two\-state mean prediction lands inside the unseen prompts’ bootstrap95%95\\%CI in both directions,

5\.14∈\[4\.69,6\.02\],5\.35∈\[4\.53,5\.77\],5\.14\\in\[4\.69,6\.02\],\\qquad 5\.35\\in\[4\.53,5\.77\],but so does the geometric’s: on these small folds the mean is a weak discriminator, which is why the falsifiable claim is distribution\-level\.

The absolute rejection of the full\-sample two\-state fit is the finite\-sample face of the noise ceiling \(Proposition[2](https://arxiv.org/html/2609.00355#Thmproposition2)\)\. ∎

### B\.5Entropy\-Keyed Width Allocation

The economics of width for one\-pass block drafting are cited machinery, not ours: the node\-sum value of a tree, top\-prefix optimality, the marginal stopping rule, and shared\-budget global best\-first allocation are taken as given from prior work\([Chen et al\. 2026](https://arxiv.org/html/2609.00355#bib.bib9);[Ringel and Romano 2026](https://arxiv.org/html/2609.00355#bib.bib10);[Zhang et al\. 2026b](https://arxiv.org/html/2609.00355#bib.bib11)\)\. Our addition is the*content key*: the law indexes a round’s node masses by its entropy, so the round’s marginal value of width tracks the remaining miss\-mass

V⁡\(H\)∝𝔼⁡\[a∣H\]⋅\(1−pm​\(H\)\),V\(H\)\\;\\propto\\;\\mathbb\{E\}\[a\\mid H\]\\cdot\\big\(1\-p\_\{\\mathrm\{m\}\}\(H\)\\big\),\(5\)small when acceptance is hopeless or already deep, and peaking at intermediate entropy\.

###### Proposition 4\(Optimal content\-indexed allocation\)\.

Across rounds with entropies\{Hr\}\\\{H\_\{r\}\\\}under a shared verifier budget, the allocation maximizing total expected committed length is global best\-first with entropy\-indexed masses: marginal width goes to the rounds of highestV⁡\(Hr\)V\(H\_\{r\}\), that is, the intermediate\-entropy band, until the marginal mass falls below the token\-time exchange rate\.

###### Proposition 5\(Small headroom, and when\)\.

The gainΔadapt\\Delta\_\{\\mathrm\{adapt\}\}of the optimal content\-indexed policy over uniform width at equal budget is positive iff width is genuinely scarce and non\-trivial mass sits in the intermediate\-entropy band, and it shrinks to zero as the entropy distribution becomes bimodal, and a largeb1b\_\{1\}thins the band \(width∼1/b1\\sim 1/b\_\{1\}\) and drivesΔadapt→0\\Delta\_\{\\mathrm\{adapt\}\}\\to 0\.

###### Proposition 6\(Regret of the entropy\-keyed policy\)\.

LetΠH\\Pi\_\{H\}allocate byV^​\(H\)\\widehat\{V\}\(H\)from the fitted law andΠ⋆\\Pi^\{\\star\}by realized round values\. ThenRegret⁡\(ΠH\)≤∑r∈Π⋆​△​ΠH\|Vr⋆−V^​\(Hr\)\|≤2​K​𝔼​\|V⋆−V^​\(H\)\|\\mathrm\{Regret\}\(\\Pi\_\{H\}\)\\leq\\sum\_\{r\\in\\Pi^\{\\star\}\\triangle\\Pi\_\{H\}\}\|V\_\{r\}^\{\\star\}\-\\widehat\{V\}\(H\_\{r\}\)\|\\leq 2K\\,\\mathbb\{E\}\|V^\{\\star\}\-\\widehat\{V\}\(H\)\|, and noσ⁡\(H\)\\sigma\(H\)\-measurable policy drives the regret below the irreducibleHH\-conditional value variance \(Proposition[2](https://arxiv.org/html/2609.00355#Thmproposition2)\)\.

###### Proof of Propositions[4](https://arxiv.org/html/2609.00355#Thmproposition4)to[6](https://arxiv.org/html/2609.00355#Thmproposition6)\.

*Proposition[4](https://arxiv.org/html/2609.00355#Thmproposition4)\.*By the cited node\-sum decomposition, the shared\-budget maximizer is the top\-KKglobally highest\-mass nodes\. Substituting entropy\-indexed masses and differentiating the truncated\-geometric node sum gives

V⁡\(H\)∝𝔼⁡\[a∣H\]​\(1−pm​\(H\)\),V\(H\)\\propto\\mathbb\{E\}\[a\\mid H\]\\,\(1\-p\_\{\\mathrm\{m\}\}\(H\)\),so best\-first is best\-VV\-first with the cited stopping rule keyed byHH\.

*Proposition[5](https://arxiv.org/html/2609.00355#Thmproposition5)\.*Three regimes:

- •If width is affordable everywhere, the stopping rule already saturates every round and the budget does not bind\.
- •If almost all rounds havepm≈1p\_\{\\mathrm\{m\}\}\\approx 1or𝔼⁡\[a∣H\]≈0\\mathbb\{E\}\[a\\mid H\]\\approx 0, thenV≈0V\\approx 0and reallocating moveso⁡\(1\)o\(1\)\.
- •When the budget binds andVVdisperses, moving a slot from a low\-VVto a high\-VVround strictly gains, and by rearrangement the gain scales with the concentration ofVV, which bimodality and band\-thinning kill\.

*Proposition[6](https://arxiv.org/html/2609.00355#Thmproposition6)\.*A standard exchange argument:ΠH\\Pi\_\{H\}maximizes∑rV^r​Πr\\sum\_\{r\}\\widehat\{V\}\_\{r\}\\Pi\_\{r\}, so the regret telescopes onto the disagreement set weighted by value\-estimation errors, and conditional\-mean orthogonality bounds anyσ⁡\(H\)\\sigma\(H\)\-measurable value estimate from below by the conditional value variance\. ∎

Empirically the headroom is as small as Proposition[5](https://arxiv.org/html/2609.00355#Thmproposition5)predicts\. At equal verifier budget the oracle width policy improves accepted value over the best fixed width by\+8\.5%\+8\.5\\%/\+3\.8%\+3\.8\\%\(captioning/TextVQA\) at budget3131and\+3\.1%\+3\.1\\%/\+4\.3%\+4\.3\\%at4747\(Figure[5](https://arxiv.org/html/2609.00355#A2.F5)\)\.

We therefore run fixed width in all main experiments and report the controller analysis as a characterization, not a headline lever\.

Figure 5:Accepted\-value versus budget Pareto for oracle, entropy\-keyed, and fixed\-width policies\.

## Appendix CClassic Two\-Model Speculative Decoding

Classic two\-model speculation, the original method that head\-based drafting later specialized, runs on our target through the official HF assisted\-generation path with three draft models: image\-conditioned Qwen3\-VL\-4B and Qwen3\-VL\-2B, and a text\-only Qwen3\-1\.7B in the spirit of the strong baseline of the first multimodal study\([Gagrani et al\. 2024](https://arxiv.org/html/2609.00355#bib.bib23)\)\(Table[4](https://arxiv.org/html/2609.00355#A3.T4)\)\.

The44B draft accepts long blocks \(τ\\tau3\.533\.53to4\.474\.47, rising with grounding\) yet decodes at0\.610\.61to0\.80×0\.80\\times, since a half\-size draft costs roughly half a target forward for each drafted token and no acceptance amortizes that\. The22B draft accepts less \(3\.113\.11to3\.973\.97\) and stays below1×1\\timesthroughout \(0\.660\.66to0\.90×0\.90\\times\), so halving the draft again narrows the deficit without closing it\.

The text\-only draft isolates what vision access is worth: acceptance collapses to1\.421\.42to2\.122\.12while the least grounded tasks still accept the least, so on grounded VLM workloads an image\-conditioned drafter is not an optional refinement but the reason drafting pays\.

Table 4:Full draft\-level breakdown behind the Classic SD rows of Table[2](https://arxiv.org/html/2609.00355#S5.T2), which are the44B and text\-only blocks\. All three drafts run the same protocol on the same prompts, so the blocks are read against one another\. Speedup is the within\-engine ratio of the same run\.
## Appendix DImplementation Details

#### Primary drafter\.

The primary drafter is a block head in the DFlash family\. It reads the target’s fused hidden state and proposes the whole block in one forward pass, with the vision tower and target decoder frozen\. The final head is selected by held\-out top\-11accuracy at each offset\. Table[5](https://arxiv.org/html/2609.00355#A4.T5)lists its configuration, with every value copied from the training run\.

Table 5:The primary Qwen3\-VL\-8B drafter training card\.Remaining constants of the run: intermediate size1228812288, block horizonL=15L\{=\}15, image long side768768px, AdamW betas0\.90\.9/0\.950\.95with no weight decay and gradient clipping1\.01\.0, seed00, one A6000\-class GPU, bf16 weights \(2\.12\.1GB\)\.

#### OCR\-retrained drafter\.

Replaces the0\.8offset0\.8^\{\\text\{offset\}\}decay with a survival\-weighted coverage loss and adds10\.210\.2K self\-generated OCR rows \(DocVQA, ChartQA, InfoVQA,26\.226\.2K total\), warm\-started from the primary head\. The checkpoint used is a partial\-epoch \(40%40\\%\) one, sufficient to show the selective grounded\-task lift of Table[17](https://arxiv.org/html/2609.00355#A6.T17)\.

#### Second and third targets\.

The Qwen3\-VL\-4B adapter follows the primary recipe, passing the fp32 lossless gate and reproducing the transfer trends in Section[6](https://arxiv.org/html/2609.00355#S6.SS0.SSS0.Px2)\. For Qwen2\.5\-VL\-7B \(Table[2](https://arxiv.org/html/2609.00355#S5.T2)\), the head is trained from scratch on26,18026\{,\}180target greedy generations \(80008000caption,80008000TextVQA, and10,18010\{,\}180disjoint OCR rows\) for seven epochs at global batch4848\. Evaluated on an A100\-80GB with a budget\-6363tree, it matches baseline prompt counts task by task \(≈47,800\{\\approx\}47\{,\}800tokens over≈12,400\{\\approx\}12\{,\}400rounds\) with zero fp32 divergences\. An H100 crosscheck replicatesτ\\tauwithin2%2\\%\(4\.184\.18/3\.213\.21/3\.153\.15\); we report A100 figures as the primary benchmark to avoid batch\-11kernel launch latency distortions on H100\. The Llama\-3\.2\-11B\-Vision adapter is warm\-started from the Llama\-3\.1\-8B DFlash head \(held\-out top\-11@11reaches0\.780\.78\), where cross\-attention visual injection accounts for the boundary behavior in Section[6](https://arxiv.org/html/2609.00355#S6.SS0.SSS0.Px3)\.

#### Matched\-training architecture comparison\.

Both architectures are trained under identical conditions: the same26,19626\{,\}196\-row OCR corpus, frozen target, global batch2424, three epochs, and SpecForge\. On the256256\-prompt held\-out probe \(greedy,6464tokens\), our budget\-6363tree achieves an acceptance ratio of2\.532\.53across three tasks and2\.732\.73\(4\.374\.37vs\.1\.601\.60\) over all five domains against the depth\-33EAGLE3 chain \(depth\-77adds≤0\.01\\leq 0\.01\)\. This advantage is consistent across probe sizes \(1616–256256prompts\) and holds when retrained on ALLaVA\-Instruct\([Chen et al\. 2024a](https://arxiv.org/html/2609.00355#bib.bib47)\)\(2\.042\.04vs\.1\.291\.29\)\. In a timed decode\-only A100 harness \(2020prompts/task,128128tokens\), GLANCE delivers2\.362\.36–2\.59×2\.59\\timesdecode speedup over autoregressive decoding versus1\.131\.13–1\.17×1\.17\\timesfor EAGLE3 \(acceptance ratio2\.242\.24, decode\-time ratio2\.122\.12\), confirming a robust2\.22\.2–2\.7×2\.7\\timesarchitectural advantage\. These standalone probe figures are decode\-only in HF; deployed system metrics and sampling results appear in Tables[2](https://arxiv.org/html/2609.00355#S5.T2)and[3](https://arxiv.org/html/2609.00355#S5.T3)\.

#### Acceptance protocol\.

Common decode settings: batch size11,T=0T\{=\}0, maximum256256new tokens,2020prompts per domain \(warmup excluded\)\.

- •Engines & metric:EAGLE3\-VL \(production AQ\-MedAI\) and n\-gram baselines run on vLLM0\.22\.10\.22\.1; ours uses an HF tree harness\. Speedup is the ratio of mean per\-token decode time against the autoregressive baseline, aligning with Medusa’s step\-overhead formulation\([Cai et al\. 2024](https://arxiv.org/html/2609.00355#bib.bib7)\)\. We report end\-to\-end wall\-clock speedup rather than theoreticalτ/forward\\tau/\\text\{forward\}metrics to account for round\-level overhead\.
- •Output equivalence \(Table[7](https://arxiv.org/html/2609.00355#A4.T7)\):Evaluated over6060prompts \(2020/task\)\. Under fp32, GLANCE achieves60/6060/60exact token matches against target greedy decoding, matching the drafter\-free reference\. In contrast, ViSpec and Medusa yield0/210/21exact matches across official, deep, and light tree configurations\. Note that bf16 exhibits numerical drift even in drafter\-free decoding \(25/6025/60exact matches\)\.
- •Baselines & ablations: - –Prompt lookup:Evaluated via vLLM; acceptance length reflects proposed block length rather than step\-levelτ\\tau, with wall\-clock speedup unaffected\. - –ViSpec & Medusa:Home\-target ViSpec uses its default3030\-token tree \(33passes/round\) withτ\\taualigned by adding the committed token\. On Qwen3\-VL\-8B, both heads follow official two\-stage training and length\-controlled evaluation to prevent early\-termination artifacts\. - –Ablations:Tree\-versus\-chain gains are isolated in Table[9](https://arxiv.org/html/2609.00355#A5.T9)\. Classic SD endpoints reflect the three\-draft sweep in Table[4](https://arxiv.org/html/2609.00355#A3.T4)\. Mismatched\-image probes use6060prompts/domain on a shared trajectory with prompt\-clustered bootstrap CIs\.

We include n\-gram lookup as the training\-free baseline; lookahead decoding is omitted due to a lack of multimodal target support\.

Table 6:EAGLE3\-VL chain\-length sweep in vLLM, the only verifier\-budget knob that engine exposes for this head\.systemarithmeticidentical*Same precision on both arms, our harness*autoregressive, no drafterbf1625/6025/60autoregressive, no drafterfp3260/6060/60n\-gram lookup \(PLD\)‡bf1627/6027/60GLANCE, budget\-6363treebf1627/6027/60GLANCE, budget\-6363treefp32𝟔𝟎/𝟔𝟎\\mathbf\{60/60\}GLANCE, matched\-training headbf1633/6033/60EAGLE3 head, matched trainingbf1619/6019/60*Same precision, each system in its released harness*EAGLE3\-VL \(production, SGLang\)bf1648/6048/60ViSpec \(official recipe\)bf160/630/63Medusa \(same codebase\)bf160/630/63ViSpec \(released head, home target\)bf1699/52099/520*Cross\-precision reference*autoregressive, no drafterbf16 vs fp3217/6017/60

Table 7:Output equivalence behind the lossless column of Table[2](https://arxiv.org/html/2609.00355#S5.T2)\.*Arithmetic*: the precision both arms run in, except the last block, which scores bf16 against an fp32 reference\.‡\\ddagger: measured on an A6000 while the rest of that block ran on the Blackwell card, and a bf16 identity rate depends on which kernels the card selects, so this row is read against1\.01\.0and not against its neighbours\.
#### Timing protocol\.

All wall\-clock numbers are decode\-only, greedy, batch one, and byte\-identity\-gated\. Absolute ms/token is never compared across engines, and only within\-engine ratios are reported\. An eager HF loop is roughly1717times slower in absolute terms than the native runner, which is exactly why cross\-engine absolutes are meaningless\.

- •Hardware: the Qwen3\-VL\-8B rows of Table[2](https://arxiv.org/html/2609.00355#S5.T2), ours and the engine baselines alike, run on one RTX PRO 6000 Blackwell, both arms of every ratio on that card, the home\-target rows on A100\-80GB, the native full\-loop on A6000 \(the88K bucket re\-measured on A100\-80GB, all systems on one card\), the HF tree harness, including the budget sweep, on H100 NVL, and the HF forward\-level rows and the matched\-training timed harness on A100\. Only within\-bucket ratios are reported, so each comparison shares one card\. The separation is measurable: replaying the identical home\-target run at matched prompt counts on an A100 and on an H100 movesτ\\tauby at most2\.0%2\.0\\%but moves the speedup by−0\.9\-0\.9to\+8\.9%\+8\.9\\%, because autoregressive decoding is bandwidth\-bound while our loop is not, so the two arms rescale differently across cards\. Table[2](https://arxiv.org/html/2609.00355#S5.T2)reports the lower, A100 figures\.
- •Same\-engine EAGLE3\-VL: both methods run in the same production engine, both CUDA\-graph captured, with an eager cross\-check\.
- •Statistics: prompt counts differ by condition and are stated with each measurement, from101101in each task of Table[1](https://arxiv.org/html/2609.00355#S5.T1)to200200on the home target\. All reported ratios carry bootstrap95%95\\%confidence intervals, listed in Appendix[E](https://arxiv.org/html/2609.00355#A5)\. In the same\-engine run the autoregressive arm lands within1%1\\%of itself across all five tasks, so the workloads differ in what they generate, not in what they cost to decode\. Three replicates of the22K\-context same\-engine comparison reproduce its margin\.

## Appendix ETiming Ledger

#### Cost distribution\.

Across the five tasks, visual encoder prefill accounts for only33–14%14\\%of a token’s end\-to\-end cost at the resolutions and lengths in Table[1](https://arxiv.org/html/2609.00355#S5.T1), with the decode loop dominating the remainder\. This prefill share decreases further as generated outputs grow\.

#### Bootstrap intervals \(Table[1](https://arxiv.org/html/2609.00355#S5.T1)\)\.

The margin column \(EAGLE3\-VLms/GLANCEms−1\\text\{EAGLE3\-VL\}\_\{\\text\{ms\}\}/\\text\{GLANCE\}\_\{\\text\{ms\}\}\-1\) matches the printed speedup ratios\. Paired bootstrap over the101101shared prompts \(80008000resamples\) yields: captioning−16\.2%\-16\.2\\%\[−17\.4,−15\.0\]\[\-17\.4,\-15\.0\]\(1/1011/101wins\), TextVQA−19\.1%\-19\.1\\%\[−20\.7,−17\.4\]\[\-20\.7,\-17\.4\]\(1/1011/101wins\), InfographicVQA\+7\.6%\+7\.6\\%\[\+5\.3,\+9\.7\]\[\+5\.3,\+9\.7\]\(79/10179/101wins\), DocVQA\+5\.4%\+5\.4\\%\[\+2\.6,\+8\.3\]\[\+2\.6,\+8\.3\]\(71/10171/101wins\), and ChartQA\+6\.0%\+6\.0\\%\[\+4\.0,\+8\.1\]\[\+4\.0,\+8\.1\]\(70/10170/101wins\)\. All intervals exclude zero, confirming consistent prompt\-level outcomes rather than outlier effects\. Geometric means are\+6\.3%\+6\.3\\%across the three grounded tasks,−17\.7%\-17\.7\\%over the two open tasks, and−4\.0%\-4\.0\\%overall\. Averaging prompt\-level ratios instead of taking the ratio of means shifts figures by≤1\.1\\leq 1\.1points without reordering\.

#### Key ledger observations\.

Table[8](https://arxiv.org/html/2609.00355#A5.T8)details wall\-clock benchmarks across engines, context lengths, and samples; Table[10](https://arxiv.org/html/2609.00355#A5.T10)presents the verifier\-budget saturation sweep; Table[9](https://arxiv.org/html/2609.00355#A5.T9)reports by\-domain forward\-level wall\-clock; and Table[2](https://arxiv.org/html/2609.00355#S5.T2)provides domain\-level acceptance margins\. Three findings stand out:

1. 1\.Value of tree width:The zero\-confound contrast \(identical head run as tree vs\. chain in the same harness\) yields1\.36×1\.36\\times\.
2. 2\.Conservative reporting:When timed under identical harness conditions with EAGLE3\-VL at its measured acceptance, grounded tasks yield1\.661\.66/1\.761\.76/1\.99×1\.99\\times\(geomean1\.80×1\.80\\times\)\. We conservatively report the1\.26×1\.26\\timesprojection, which grants EAGLE3\-VL its production acceptance rate\.
3. 3\.Long\-context robustness:When scaling context from22K to∼6\.7\{\\sim\}6\.7K tokens, acceptance drops3\.94→3\.393\.94\\to 3\.39\(13\.9%13\.9\\%\) for GLANCE versus3\.87→3\.533\.87\\to 3\.53\(8\.7%8\.7\\%\) for EAGLE3\-VL\. However, because GLANCE requires only11draft pass per round versus88for EAGLE3\-VL, end\-to\-end decode times remain within2\.3%2\.3\\%of each other despite the steeper acceptance decline\.

Table 8:Wall\-clock comparisons by engine, mode, and context length, decode\-only and byte\-identity\-gated\. Percentages are\(EAGLE3\-VLms−GLANCEms\)/EAGLE3\-VLms\(\\text\{EAGLE3\-VL\}\_\{\\text\{ms\}\}\-\\text\{GLANCE\}\_\{\\text\{ms\}\}\)/\\text\{EAGLE3\-VL\}\_\{\\text\{ms\}\}, a ratio of means over the shared prompts, signed so that positive means GLANCE is faster, with a paired bootstrap interval of that same statistic\.Table 9:Domain\-level ms/token behind the chain\-row speedup column of Table[2](https://arxiv.org/html/2609.00355#S5.T2)\. The geomean row carries the tree\-versus\-autoregression and zero\-confound tree\-versus\-chain contrasts cited in Section[5\.4](https://arxiv.org/html/2609.00355#S5.SS4)\. The EAGLE3\-VL column is our faithful HF re\-implementation, whose acceptance falls short of the production vLLM head’s\. Dividing the printed ms/token would therefore flatter us, so the*vs\. EAGLE3\-VL*column instead charges EAGLE3\-VL its production acceptance at the same round\-level cost, which is the conservative comparison and the one we report\.Table 10:Verifier\-budget sweep on all five tasks in the same\-engine tree harness\. All four budgets of a task run in one process and share that task’s autoregressive baseline, so the speedup column is a function ofNNalone\.Table 11:Forward\-level cost of one decode round on Qwen3\-VL\-8B, all three components measured in one process on one A6000, decode\-only\. The packed tree verify and the one\-pass draft replay from CUDA graphs\.The same asymmetry is visible in the production head’s own sweep, this time inside SGLang, where the engine exposes the draft\-step count directly rather than the single chain\-length knob vLLM offers in Table[6](https://arxiv.org/html/2609.00355#A4.T6)\. Regressing its round time on the number of draft passes over five configurations \(44to4848draft tokens at33to1010steps, captioning, one A6000\) gives24\.824\.8ms plus1\.821\.82ms for every sequential draft pass, with a constant term worth1\.031\.03target forwards\. At batch one the verify is therefore near\-flat in tree width while depth is paid a pass at a time\. Widening our tree4\.2×4\.2\\timescosts33to5%5\\%of a round, whereas taking that head from two passes to eight costs38%38\\%of one, with the component costs in Table[11](https://arxiv.org/html/2609.00355#A5.T11)\.

## Appendix FAdditional Transfer and Robustness Results

#### Statistics behind the characterization\.

Table[13](https://arxiv.org/html/2609.00355#A6.T13)gives the task\-level correlations and decile gaps behind Section[6](https://arxiv.org/html/2609.00355#S6): entropy correlates with acceptance in every task while the image\-saliency signalvtv\_\{t\}does not survive conditioning on entropy \(partial\|r\|≤0\.09\|r\|\\leq 0\.09, within\-decile gaps≈1\\approx 1\), which Figure[7](https://arxiv.org/html/2609.00355#A6.F7)draws\. Figure[8](https://arxiv.org/html/2609.00355#A6.F8)shows the task\-level law fits behind Table[12](https://arxiv.org/html/2609.00355#A6.T12)\.

#### The law on the two remaining tasks\.

Round logs collected after the main fits close the five\-task family, and Table[12](https://arxiv.org/html/2609.00355#A6.T12)carries their fits alongside the original three\. InfoVQA follows the same streaming protocol as the originally logged tasks, while ChartQA is logged under the paper’s templated evaluation prompts and carries the steepest slope of the five\. The isotonic curveR2R^\{2\}stays high on both,0\.9790\.979and0\.9910\.991, so the monotone claim does not lean on the closed form\. Both floors clear their own caps,4\.424\.42against3\.223\.22and6\.926\.92against4\.504\.50, extending the copy\-regime excess of Theorem[2](https://arxiv.org/html/2609.00355#Thmtheorem2)beyond DocVQA, andP⁡\(a≥8\)P\(a\{\\geq\}8\)rises monotonically with grounding,0\.00\.0,0\.90\.9,2\.62\.6,4\.44\.4and11\.5%11\.5\\%in the task order of Table[16](https://arxiv.org/html/2609.00355#A6.T16)\.

Table 12:Entropy\-acceptance fits behind Section[6](https://arxiv.org/html/2609.00355#S6): affine\-logit coefficients, curve and roundR2R^\{2\}, ceiling fraction, and meanaa\.Figure 6:Entropy\-acceptance fits for three targets \(a\) and task\-level mean accepted length on Qwen2\.5\-VL\-7B \(b\)\.Figure 7:Acceptance split by image saliency within entropy deciles\.Table 13:Correlations and decile gaps\. Hereaais accepted length,HHentropy,vtv\_\{t\}image saliency\.Figure 8:Task\-level entropy\-acceptance fits behind Table[12](https://arxiv.org/html/2609.00355#A6.T12), ChartQA under its templated prompts\.
#### Causal\-probe statistics\.

Behind the mismatched\-image probe of Section[6](https://arxiv.org/html/2609.00355#S6.SS0.SSS0.Px1)\(protocol in Appendix[D](https://arxiv.org/html/2609.00355#A4)\)\. The correct image lengthens accepted blocks, every interval excluding zero and the effect scaling with grounding:

- •captioning:\+3\.5%\+3\.5\\%\(95%95\\%CI\[\+1\.5,\+5\.7\]\[\+1\.5,\+5\.7\]\),
- •TextVQA:\+6\.8%\+6\.8\\%\(95%95\\%CI\[\+4\.2,\+9\.4\]\[\+4\.2,\+9\.4\]\),
- •DocVQA:\+11\.7%\+11\.7\\%\(95%95\\%CI\[\+7\.7,\+16\.3\]\[\+7\.7,\+16\.3\]\)\.

Matching images sharply reduces target mean next\-token entropy; controlling for entropy, the residual visual effect is non\-positive \(pooled coefficient−0\.07\-0\.07\)\.

#### Cross\-attention boundary statistics\.

In Section[6](https://arxiv.org/html/2609.00355#S6.SS0.SSS0.Px3), Llama\-3\.2\-11B\-Vision accepts tokens in inverse grounding order: captioning \(3\.763\.76\), DocVQA \(3\.093\.09\), and TextVQA \(2\.762\.76\) in an fp32 run \(1515prompts/task,739739rounds\)\. Entropy fits here and in Table[14](https://arxiv.org/html/2609.00355#A6.T14)stem from a separate bf16 run \(2525prompts,13151315rounds, captioningτ=3\.86\\tau\{=\}3\.86\), making the two non\-interchangeable\. Thus, the scaling law holds within captioning but inverts sign on grounded short\-answer domains\.

#### Law transfer, both axes\.

Table[14](https://arxiv.org/html/2609.00355#A6.T14)details the transfer axes from Section[6](https://arxiv.org/html/2609.00355#S6.SS0.SSS0.Px2)\. In Panel A, EAGLE3\-VL uses our HF re\-implementation solely to test architectural form/sign transfer \(distinct from Table[2](https://arxiv.org/html/2609.00355#S5.T2)\), while the degenerate n\-gram fit marks the no\-model boundary\. In Panel B, only Qwen3\-VL\-8B uses our block\-diffusion adapter \(other targets use same\-family speculative proxies\), confirming the empirical*form*is target\-intrinsic rather than drafter\-specific\. Excluding a degenerate Qwen2\.5\-VL EAGLE3 baseline, all five pre\-registered conditions pass\.

Panel A: varied drafter \(same target Qwen3\-VL\-8B\) Panel B: varied target \(each non\-adapter row a same\-family proxy\)

Table 14:Law transfer: Panel A varies the drafter on the fixed Qwen3\-VL\-8B target \(its first row is also the reference row for Panel B\)\. Panel B varies the target, and Figure[6](https://arxiv.org/html/2609.00355#A6.F6)plots the same fits\. Meanaais the accepted length without the bonus token in both panels\.
#### Beyond vision\.

The cross\-modality fits of Table[15](https://arxiv.org/html/2609.00355#A6.T15)span speech recognition, code, chat, and a non\-Qwen backbone, all withb1\>0b\_\{1\}\>0\. The mechanism differs by modality in exactly the way the two\-state analysis of Appendix[B\.4](https://arxiv.org/html/2609.00355#A2.SS4)anticipates: grounded VL and code sit on the determinate\-copy route \(long verbatim runs\), while speech behaves like a uniform position\-level lift, and both are limits of the same survival family\.

Table 15:Functional\-form transfer across modalities and to a non\-Qwen backbone\.
#### Head size, matched\.

Behind Figure[9](https://arxiv.org/html/2609.00355#A6.F9): our head has1\.051\.05B parameters against EAGLE3\-VL’s11\-layer399\.7399\.7M \(0\.800\.80GB\)\. The capacity conclusion is drawn in Section[5\.5](https://arxiv.org/html/2609.00355#S5.SS5)\.

A like\-for\-like width match is not expressible for an autoregressive head, whose depth costs sequential passes, which is precisely the asymmetry of Section[3\.2](https://arxiv.org/html/2609.00355#S3.SS2)\.

Figure 9:Acceptance against verifier budget for the identical head run as chain and as tree \(a\), and draft cost against chain depth or budget \(b\)\.
#### Conditioning ablation and n\-gram floor\.

Table[16](https://arxiv.org/html/2609.00355#A6.T16)tabulates the conditioning ablation cited in Section[5\.5](https://arxiv.org/html/2609.00355#S5.SS5)\. Pooled over the five tasks and their300300prompts, the direct\-visual\-fusion lift is\+14\.1%\+14\.1\\%\(95%95\\%CI\[\+12\.8,\+15\.5\]%\[\+12\.8,\+15\.5\]\\%\), and it is the grounded tasks that carry it\. The model\-free n\-gram lookup floor is tabulated in Table[2](https://arxiv.org/html/2609.00355#S5.T2)\. It stays below the wide tree on every task\.

Table 16:Conditioning ablation on Qwen3\-VL\-8B: acceptance lengthτ\\tauat budget3131, sixty prompts in each task\. The middle column replaces the target’s fused states with the hidden states of a text\-only Qwen3\-8B at the text positions and zeros at the image positions, so that drafter never sees the image\.Table 17:OCR\-targeted retraining on Qwen3\-VL\-8B: acceptance lengthτ\\taubefore and after adding OCR\-heavy training data, at budgets3131and6363

Similar Articles

Speculative Decoding with a Speculative Vocabulary

arXiv cs.CL

This paper proposes SpecVocab, a method for selecting a per-step vocabulary subset for the draft model in speculative decoding, achieving higher acceptance length and up to 8.1% throughput improvement over EAGLE-3.

What is Speculative Decoding? (trending on paperswithco.de) [R]

Reddit r/MachineLearning

Speculative decoding is an inference optimization technique that uses a fast draft model to propose future tokens verified in parallel by a larger model, improving LLM generation speed. The article highlights its trending status on Papers with Code and a recent SGLang blog post about state-of-the-art latencies using DFlash models.