speculative-decoding

Tag

Cards List
#speculative-decoding

TSS: Target-Side Sparsification for Speculative Decoding in Domain-Specific Large Language Models

arXiv cs.CL ↗ · 2026-09-23 Cached

The paper proposes TSS, a target-side sparsification framework for speculative decoding in domain-specific LLMs, which improves inference efficiency and performance by skipping selected target layers.

0 favorites 0 likes
#speculative-decoding

Fork of FreeToken with DeepSeek-V4.1, vision and speculative decoding (2x3090 numbers inside)

Reddit r/LocalLLaMA ↗ · 2026-09-22

A developer shares a fork of FreeToken, an edge-native MoE serving engine, with added support for DeepSeek-V4.1, vision capabilities for Qwen models, and speculative decoding, including benchmarks on dual RTX 3090 hardware.

0 favorites 0 likes
#speculative-decoding

The Limits of Speculation: Bounding Speculative Decoding in Mixture-of-Experts

arXiv cs.LG ↗ · 2026-09-22 Cached

This paper formulates speculative decoding in Mixture-of-Experts models as a Stochastic Shortest Path problem and uses a diagnostic Oracle to demonstrate that optimal decisions follow a necessary condition balancing cost and progress.

0 favorites 0 likes
#speculative-decoding

TreeSpark: Calibrated, Load-Adaptive Draft Trees for Semi-Autoregressive Speculative Decoding

arXiv cs.CL ↗ · 2026-09-22 Cached

TreeSpark introduces calibrated, load-adaptive draft trees for semi-autoregressive speculative decoding in language models, increasing draft token acceptance by 15-25% and speeding up decoding by 8-14%.

0 favorites 0 likes
#speculative-decoding

[Splash Engine] Qwen3.8-27B in native 8-bit at 37–55 tok/s on Apple Silicon: Extending Splash to Q8, 256k context scaling, and the "Reasoning Cliff"

Reddit r/LocalLLaMA ↗ · 2026-09-21

Extended Splash Engine to support native 8-bit Qwen3.8-27B on Apple Silicon, achieving 37-55 tok/s without quantization degradation and scaling up to 256k context.

0 favorites 0 likes
#speculative-decoding

To Copy or Not to Copy: Controlling Speculative Decoding via Intrinsic Model Signals

arXiv cs.CL ↗ · 2026-09-18 Cached

SwitchSD is an adaptive framework for speculative decoding in LLMs that uses intrinsic model signals to dynamically switch between neural drafting and context-based copying, achieving up to 15% throughput gains over baselines like EAGLE3.

0 favorites 0 likes
#speculative-decoding

Zarya: A Hybrid Autoregressive--Masked Diffusion Language Model with Flexible Training and Dual-Mode Inference

arXiv cs.CL ↗ · 2026-09-18 Cached

Zarya is a hybrid language model that jointly optimizes autoregressive and masked diffusion objectives for flexible training and dual-mode inference, with publicly released models in sizes 0.6B, 1.7B, and 4B.

0 favorites 0 likes
#speculative-decoding

Shapelearn Qwen 3.8 27B (13.1 GB VRAM)

Hacker News Top ↗ · 2026-09-18 Cached

ByteShape releases full ShapeLearn quantized versions of the Qwen 3.8 27B model in GGUF format, with benchmarking showing improvements in quality-speed frontier and support for speculative decoding.

0 favorites 0 likes
#speculative-decoding

@ngxson: My talk about llama.cpp speculative decoding (MTP, dflash, dspark) at dotAI replay available soon!

X AI KOLs Following ↗ · 2026-09-17 Cached

A talk about llama.cpp speculative decoding methods (MTP, dflash, dspark) at the dotAI conference will have its replay available soon.

0 favorites 0 likes
#speculative-decoding

Intel releases OpenVINO 2026.4

Reddit r/LocalLLaMA ↗ · 2026-09-17

Intel releases OpenVINO 2026.4 with expanded AI model support, performance enhancements like multi-token prediction, and new features for profiling and inference across CPUs, GPUs, and NPUs.

0 favorites 0 likes
#speculative-decoding

The Inference Engineering Pareto Atlas: Which Optimizations Dominate the Cost, Quality, and Latency Frontier?

arXiv cs.AI ↗ · 2026-09-17 Cached

This paper constructs a cost-quality-latency Pareto atlas for LLM inference optimizations, using a calibrated simulator to evaluate configurations and combinations across different hardware and regimes.

0 favorites 0 likes
#speculative-decoding

FlexEE: Self-Speculative and KV-Compatible Early Exiting for Offloading-Aware LLM Inference

arXiv cs.AI ↗ · 2026-09-16 Cached

FlexEE introduces a self-speculative and KV-cache-compatible early exiting framework for efficient LLM inference in offloading deployments, achieving significant speedups on Llama models with minimal accuracy degradation.

0 favorites 0 likes
#speculative-decoding

DeepSeek V4.1F Q4 on M3 Ultra with native DSpark MTP (40tps / 800tps)

Reddit r/LocalLLaMA ↗ · 2026-09-15

The author optimized DeepSeek V4.1 Flash for Apple M3 Ultra, achieving up to 40 t/s decode speed with DSpark speculative decoding while maintaining byte-identical accuracy to the upstream model.

0 favorites 0 likes
#speculative-decoding

R9V Update: now ~100 tok/s in TG on Qwen3.8 Flash Next IQ4_XS on x2 R9700 + 128GB RAM. Fixed crashes with n-gram SSD streaming, improved diagnostics, plus pinned images. Q4_K_XL now supported, 50 tok/s TG.

Reddit r/LocalLLaMA ↗ · 2026-09-14 Cached

R9V update delivers approximately 100 tokens per second on Qwen3.8 Flash Next with IQ4_XS on dual AMD R9700 GPUs, adds support for Q4_K_XL with 50 tok/s, fixes crashes, and enhances diagnostics.

0 favorites 0 likes
#speculative-decoding

How Lossless Is Lossless Speculative Decoding? The Role of Numerical Precision in Orthrus

Hugging Face Daily Papers ↗ · 2026-09-14 Cached

This paper investigates the losslessness of Orthrus, a hybrid autoregressive-diffusion model for inference acceleration, finding that it requires high numerical precision (FP32) for exact trajectory matching, while BF16 divergence does not impair downstream performance.

0 favorites 0 likes
#speculative-decoding

This draft model is OP on 16 GB cards for Qwen 3.8 27b

Reddit r/LocalLLaMA ↗ · 2026-09-12

A draft model for Qwen 3.8 27B is highly effective on 16 GB GPUs, achieving around 60 tokens per second on an RX 9070 XT and offering better VRAM efficiency than built-in MTP.

0 favorites 0 likes
#speculative-decoding

@TheDavidTai: We just hit 100% on https://CUDA.fast! Also look at all the different models being represented on this snapshot.

X AI KOLs Following ↗ · 2026-09-11 Cached

The Qwen 3.8 Flash Next model achieved a 100% score on the CUDA.fast benchmark with a 117.8% composite speed increase on DGX Spark, utilizing speculative decoding.

0 favorites 0 likes
#speculative-decoding

Osprey: Target-agnostic Pre-training Makes Stronger Drafters in Speculative Decoding

arXiv cs.CL ↗ · 2026-09-10 Cached

Osprey introduces a target-agnostic pre-training method for drafters in speculative decoding, improving efficiency by bootstrapping from off-the-shelf models and adapting with minimal target-specific work, achieving significant acceptance rate improvements across multiple LLMs.

0 favorites 0 likes
#speculative-decoding

X-CoSD: Communication-Efficient Cross-Vocabulary Collaborative Speculative Decoding

arXiv cs.CL ↗ · 2026-09-10 Cached

This paper introduces X-CoSD, a communication-efficient cross-vocabulary collaborative speculative decoding framework that optimizes distributed LLM inference by splitting residual resampling to reduce overhead while preserving server LLM quality.

0 favorites 0 likes
#speculative-decoding

Finally understood why my coding agent types fast on boilerplate and slow on new logic

Reddit r/AI_Agents ↗ · 2026-09-09

The author explains how speculative decoding affects coding agent speed, with higher acceptance rates on boilerplate code leading to faster typing, and discusses other factors like cache misses that impact performance.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback