speculative-decoding

Tag

Cards List
#speculative-decoding

exl3 now in ninfer-ext

Reddit r/LocalLLaMA ↗ · 9h ago Cached

ninfer-ext is an extended fork of NInfer that adds support for larger Qwen models, faster speculative decoding, agent serving features, and EXL3 quantization for Qwen3.8-27B on NVIDIA RTX 5090.

0 favorites 0 likes
#speculative-decoding

Ornith-1.5 DFlash

Reddit r/LocalLLaMA ↗ · yesterday

Ornith-1.5 models integrated with DFlash draft models for speculative decoding have been released on Hugging Face in 9B, 397B, and 35B-A3B sizes.

0 favorites 0 likes
#speculative-decoding

Mentored Decoding: Faster Inference meets Boosting

arXiv cs.LG ↗ · yesterday Cached

This paper introduces mentored decoding, a formal approach to lossy speculative decoding that improves inference speed in language models by allowing controlled divergence from the target model, connecting it to boosting theory and proving key properties for optimization.

0 favorites 0 likes
#speculative-decoding

42x Faster Prompt Lookup Drafting in llama.cpp

Reddit r/LocalLLaMA ↗ · 3d ago Cached

A blog post detailing performance optimizations in llama.cpp that make prompt lookup drafting up to 42x faster and reduce memory usage by 2.6x, based on techniques from Daniel Lemire and Martin Ankerl.

0 favorites 0 likes
#speculative-decoding

WaveFront Decoding: Parallelized Self-Speculative Decoding for Looped Language Models

Hugging Face Daily Papers ↗ · 4d ago Cached

WaveFront Decoding introduces a training-free self-speculative decoding framework for looped language models that reduces latency by concurrently batching drafting and verification, achieving up to 4.81x speedup on Huginn-3.5B.

0 favorites 0 likes
#speculative-decoding

@Anbeeld: BeeLlama v0.4.7 is out! You can now save gigabytes of VRAM when using MTP and DFlash! In the example shown in the image…

X AI KOLs Timeline ↗ · 4d ago

BeeLlama v0.4.7 is released, offering gigabytes of VRAM savings through independent drafter ubatch sizing for MTP and DFlash, along with major ROCm and CUDA performance enhancements for AI inference.

0 favorites 0 likes
#speculative-decoding

Accelerating vision-language models with LFM2.5-VL-DSpark

Hugging Face Blog ↗ · 5d ago Cached

Today, we release an experimental DSpark draft model for our vision-language model (VLM) LFM2.5-VL-3B, which adds a speculative decoding path for faster inference with minimal memory cost.

0 favorites 0 likes
#speculative-decoding

When Parallel Drafter Meets Parallel Speculative Decoding

arXiv cs.CL ↗ · 6d ago Cached

DPara is a parallel speculative decoding framework that eliminates probabilistic fallback by precomputing draft representations, achieving average speedups of 3.21× to 3.52× over autoregressive decoding on Qwen3 models.

0 favorites 0 likes
#speculative-decoding

@dzhng: Providers/labs often tune their quantization/speculative decoding strategy post-launch based on usage patterns & th…

X AI KOLs Following ↗ · 6d ago Cached

The tweet discusses how providers tune quantization and speculative decoding strategies post-launch based on usage patterns and hardware, and introduces an interactive tutorial on speculative decoding for the NeurIPS Education Track.

0 favorites 0 likes
#speculative-decoding

TSS: Target-Side Sparsification for Speculative Decoding in Domain-Specific Large Language Models

arXiv cs.CL ↗ · 2026-09-23 Cached

The paper proposes TSS, a target-side sparsification framework for speculative decoding in domain-specific LLMs, which improves inference efficiency and performance by skipping selected target layers.

0 favorites 0 likes
#speculative-decoding

Fork of FreeToken with DeepSeek-V4.1, vision and speculative decoding (2x3090 numbers inside)

Reddit r/LocalLLaMA ↗ · 2026-09-22

A developer shares a fork of FreeToken, an edge-native MoE serving engine, with added support for DeepSeek-V4.1, vision capabilities for Qwen models, and speculative decoding, including benchmarks on dual RTX 3090 hardware.

0 favorites 0 likes
#speculative-decoding

The Limits of Speculation: Bounding Speculative Decoding in Mixture-of-Experts

arXiv cs.LG ↗ · 2026-09-22 Cached

This paper formulates speculative decoding in Mixture-of-Experts models as a Stochastic Shortest Path problem and uses a diagnostic Oracle to demonstrate that optimal decisions follow a necessary condition balancing cost and progress.

0 favorites 0 likes
#speculative-decoding

TreeSpark: Calibrated, Load-Adaptive Draft Trees for Semi-Autoregressive Speculative Decoding

arXiv cs.CL ↗ · 2026-09-22 Cached

TreeSpark introduces calibrated, load-adaptive draft trees for semi-autoregressive speculative decoding in language models, increasing draft token acceptance by 15-25% and speeding up decoding by 8-14%.

0 favorites 0 likes
#speculative-decoding

[Splash Engine] Qwen3.8-27B in native 8-bit at 37–55 tok/s on Apple Silicon: Extending Splash to Q8, 256k context scaling, and the "Reasoning Cliff"

Reddit r/LocalLLaMA ↗ · 2026-09-21

Extended Splash Engine to support native 8-bit Qwen3.8-27B on Apple Silicon, achieving 37-55 tok/s without quantization degradation and scaling up to 256k context.

0 favorites 0 likes
#speculative-decoding

To Copy or Not to Copy: Controlling Speculative Decoding via Intrinsic Model Signals

arXiv cs.CL ↗ · 2026-09-18 Cached

SwitchSD is an adaptive framework for speculative decoding in LLMs that uses intrinsic model signals to dynamically switch between neural drafting and context-based copying, achieving up to 15% throughput gains over baselines like EAGLE3.

0 favorites 0 likes
#speculative-decoding

Zarya: A Hybrid Autoregressive--Masked Diffusion Language Model with Flexible Training and Dual-Mode Inference

arXiv cs.CL ↗ · 2026-09-18 Cached

Zarya is a hybrid language model that jointly optimizes autoregressive and masked diffusion objectives for flexible training and dual-mode inference, with publicly released models in sizes 0.6B, 1.7B, and 4B.

0 favorites 0 likes
#speculative-decoding

Shapelearn Qwen 3.8 27B (13.1 GB VRAM)

Hacker News Top ↗ · 2026-09-18 Cached

ByteShape releases full ShapeLearn quantized versions of the Qwen 3.8 27B model in GGUF format, with benchmarking showing improvements in quality-speed frontier and support for speculative decoding.

0 favorites 0 likes
#speculative-decoding

@ngxson: My talk about llama.cpp speculative decoding (MTP, dflash, dspark) at dotAI replay available soon!

X AI KOLs Following ↗ · 2026-09-17 Cached

A talk about llama.cpp speculative decoding methods (MTP, dflash, dspark) at the dotAI conference will have its replay available soon.

0 favorites 0 likes
#speculative-decoding

Intel releases OpenVINO 2026.4

Reddit r/LocalLLaMA ↗ · 2026-09-17

Intel releases OpenVINO 2026.4 with expanded AI model support, performance enhancements like multi-token prediction, and new features for profiling and inference across CPUs, GPUs, and NPUs.

0 favorites 0 likes
#speculative-decoding

The Inference Engineering Pareto Atlas: Which Optimizations Dominate the Cost, Quality, and Latency Frontier?

arXiv cs.AI ↗ · 2026-09-17 Cached

This paper constructs a cost-quality-latency Pareto atlas for LLM inference optimizations, using a calibrated simulator to evaluate configurations and combinations across different hardware and regimes.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback