speculative-decoding

Tag

Cards List
#speculative-decoding

@ngxson: My talk about llama.cpp speculative decoding (MTP, dflash, dspark) at dotAI replay available soon!

X AI KOLs Following ↗ · 2026-09-17 Cached

A talk about llama.cpp speculative decoding methods (MTP, dflash, dspark) at the dotAI conference will have its replay available soon.

0 favorites 0 likes
#speculative-decoding

Intel releases OpenVINO 2026.4

Reddit r/LocalLLaMA ↗ · 2026-09-17

Intel releases OpenVINO 2026.4 with expanded AI model support, performance enhancements like multi-token prediction, and new features for profiling and inference across CPUs, GPUs, and NPUs.

0 favorites 0 likes
#speculative-decoding

The Inference Engineering Pareto Atlas: Which Optimizations Dominate the Cost, Quality, and Latency Frontier?

arXiv cs.AI ↗ · 2026-09-17 Cached

This paper constructs a cost-quality-latency Pareto atlas for LLM inference optimizations, using a calibrated simulator to evaluate configurations and combinations across different hardware and regimes.

0 favorites 0 likes
#speculative-decoding

FlexEE: Self-Speculative and KV-Compatible Early Exiting for Offloading-Aware LLM Inference

arXiv cs.AI ↗ · 2026-09-16 Cached

FlexEE introduces a self-speculative and KV-cache-compatible early exiting framework for efficient LLM inference in offloading deployments, achieving significant speedups on Llama models with minimal accuracy degradation.

0 favorites 0 likes
#speculative-decoding

DeepSeek V4.1F Q4 on M3 Ultra with native DSpark MTP (40tps / 800tps)

Reddit r/LocalLLaMA ↗ · 2026-09-15

The author optimized DeepSeek V4.1 Flash for Apple M3 Ultra, achieving up to 40 t/s decode speed with DSpark speculative decoding while maintaining byte-identical accuracy to the upstream model.

0 favorites 0 likes
#speculative-decoding

R9V Update: now ~100 tok/s in TG on Qwen3.8 Flash Next IQ4_XS on x2 R9700 + 128GB RAM. Fixed crashes with n-gram SSD streaming, improved diagnostics, plus pinned images. Q4_K_XL now supported, 50 tok/s TG.

Reddit r/LocalLLaMA ↗ · 2026-09-14 Cached

R9V update delivers approximately 100 tokens per second on Qwen3.8 Flash Next with IQ4_XS on dual AMD R9700 GPUs, adds support for Q4_K_XL with 50 tok/s, fixes crashes, and enhances diagnostics.

0 favorites 0 likes
#speculative-decoding

How Lossless Is Lossless Speculative Decoding? The Role of Numerical Precision in Orthrus

Hugging Face Daily Papers ↗ · 2026-09-14 Cached

This paper investigates the losslessness of Orthrus, a hybrid autoregressive-diffusion model for inference acceleration, finding that it requires high numerical precision (FP32) for exact trajectory matching, while BF16 divergence does not impair downstream performance.

0 favorites 0 likes
#speculative-decoding

This draft model is OP on 16 GB cards for Qwen 3.8 27b

Reddit r/LocalLLaMA ↗ · 2026-09-12

A draft model for Qwen 3.8 27B is highly effective on 16 GB GPUs, achieving around 60 tokens per second on an RX 9070 XT and offering better VRAM efficiency than built-in MTP.

0 favorites 0 likes
#speculative-decoding

@TheDavidTai: We just hit 100% on https://CUDA.fast! Also look at all the different models being represented on this snapshot.

X AI KOLs Following ↗ · 2026-09-11 Cached

The Qwen 3.8 Flash Next model achieved a 100% score on the CUDA.fast benchmark with a 117.8% composite speed increase on DGX Spark, utilizing speculative decoding.

0 favorites 0 likes
#speculative-decoding

Osprey: Target-agnostic Pre-training Makes Stronger Drafters in Speculative Decoding

arXiv cs.CL ↗ · 2026-09-10 Cached

Osprey introduces a target-agnostic pre-training method for drafters in speculative decoding, improving efficiency by bootstrapping from off-the-shelf models and adapting with minimal target-specific work, achieving significant acceptance rate improvements across multiple LLMs.

0 favorites 0 likes
#speculative-decoding

X-CoSD: Communication-Efficient Cross-Vocabulary Collaborative Speculative Decoding

arXiv cs.CL ↗ · 2026-09-10 Cached

This paper introduces X-CoSD, a communication-efficient cross-vocabulary collaborative speculative decoding framework that optimizes distributed LLM inference by splitting residual resampling to reduce overhead while preserving server LLM quality.

0 favorites 0 likes
#speculative-decoding

Finally understood why my coding agent types fast on boilerplate and slow on new logic

Reddit r/AI_Agents ↗ · 2026-09-09

The author explains how speculative decoding affects coding agent speed, with higher acceptance rates on boilerplate code leading to faster typing, and discusses other factors like cache misses that impact performance.

0 favorites 0 likes
#speculative-decoding

Qwen3.8-Flash-Next in llama.cpp vs SGLang vs FreeToken: 35s vs 258s to first token at full context. My findings on new PRs coming to engines.

Reddit r/LocalLLaMA ↗ · 2026-09-08

Benchmarks comparing SGLang, llama.cpp, and FreeToken on Qwen3.8-Flash-Next at full context show SGLang achieves the fastest time to first token at 35.4s, while llama.cpp baseline takes 258.4s, with speculative decoding providing performance improvements.

0 favorites 0 likes
#speculative-decoding

Higher acceptance length, slower prose: Ling’s n=1/2/3 MTP test on one Spark

Reddit r/LocalLLaMA ↗ · 2026-09-07

Benchmark tests on Ling-3.0-flash show that higher acceptance length in multi-token prediction decreases prose throughput, with n=1 being the most efficient setting for the evaluated workloads.

0 favorites 0 likes
#speculative-decoding

Online Draft Co-Training for Speculative Decoding in Large-Scale, Long-Context RL Post-Training

Hugging Face Daily Papers ↗ · 2026-09-07 Cached

This paper presents a system to accelerate speculative decoding in large-scale RL post-training through online draft co-training, using context-parallel attention extensions and cross-stage feature transport.

0 favorites 0 likes
#speculative-decoding

Jina-OCR-v1: Efficient Document Parsing with Speculative Decoding and Dense Verifiable Rewards

arXiv cs.CL ↗ · 2026-09-04 Cached

Jina-OCR-v1 is an efficient end-to-end document parsing model that uses speculative decoding and dense verifiable rewards to achieve high accuracy and speed on low-budget GPUs, scoring 91.14 on OmniDocBench v1.6 and 83.4 on olmOCR-Bench.

0 favorites 0 likes
#speculative-decoding

Margins, Not Windows: Training-Free Per-Step Lossy Speculative Decoding

arXiv cs.CL ↗ · 2026-09-04 Cached

AdaptiveSpec is a training-free per-step speculative decoding method that adaptively adjusts token verification and draft tree shape to enhance LLM inference throughput, improving performance by up to 56% while maintaining high accuracy across benchmarks.

0 favorites 0 likes
#speculative-decoding

@PyTorch: TorchSpec is a PyTorch-native framework for training speculative decoding draft models. This release, a collaboration w…

X AI KOLs Following ↗ · 2026-09-03 Cached

TorchSpec is a PyTorch-native framework for training speculative decoding draft models, released in collaboration with vllm and demonstrated with Kimi K3 draft models on NVIDIA GB200 hardware.

0 favorites 0 likes
#speculative-decoding

@Darkolorin: Today we are releasing our speculative decoding implementation in Uzu. Biggest release since inception of our lab. Init…

X AI KOLs Timeline ↗ · 2026-09-03 Cached

Release of a speculative decoding implementation in Uzu, initially supporting Qwen3.6 27B with upcoming support for Qwen3.8 27B and Muse Glimmer.

0 favorites 0 likes
#speculative-decoding

@ViC305: I DID IT!! DeepSeek-V4-Flash-Vision EXL3 MixedK is now running VISION + DSpark speculative decoding together on ONE DGX…

X AI KOLs Timeline ↗ · 2026-09-03 Cached

User @ViC305 successfully runs DeepSeek-V4-Flash-Vision with EXL3 MixedK and DSpark speculative decoding on a single DGX Spark, achieving improved performance and fixing technical issues for multimodal AI deployment.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback