speculative-decoding

Tag

Cards List
#speculative-decoding

[Benchmark] DFlash2 vs MTP comparison. 5090RTX, Qwen 3.8 27B, Dynamic v3 GGUF, llama.cpp. Token generation, latency and available context.

Reddit r/LocalLLaMA ↗ · 2026-08-21

The benchmark compares DFlash2 and MTP techniques in llama.cpp, showing that DFlash2 offers around 20% faster token generation but reduces available context by 38%.

0 favorites 0 likes
#speculative-decoding

@TeksEdge: A localmaxxer hit ~381 tok/s on a SINGLE RTX 3090 with Qwen3.8-27B. This developer has turned a 24GB RTX 3090 into a mo…

X AI KOLs Timeline ↗ · 2026-08-21 Cached

A developer achieved up to 381 tok/s inference speed on a single RTX 3090 with the Qwen3.8-27B model using optimized techniques like DFlash2 and prefix caching, particularly effective for document-based tasks like RAG and coding assistants.

0 favorites 0 likes
#speculative-decoding

If you are wondering why Ornith 1.5 35B A3B with MTP is so slow, this is why

Reddit r/LocalLLaMA ↗ · 2026-08-20 Cached

The Ornith 1.5 35B A3B model's MTP tensors appear to be uninitialized, causing poor speculative decoding performance, and grafting the trained head from Qwen3.6-35B-A3B improves speed by 29%.

0 favorites 0 likes
#speculative-decoding

@jimmysmith1919: Another nice release today. New draft models for speculative decoding of several of our LFM2.5 models. 1.2B: https://hu…

X AI KOLs Timeline ↗ · 2026-08-20 Cached

LiquidAI releases draft models for speculative decoding to accelerate their LFM2.5 models, achieving up to 2× faster inference on H100 and Apple silicon without quality degradation.

0 favorites 0 likes
#speculative-decoding

@ramin_m_h: yesterday we made them more compressed! today we make them faster than ever with speculative decoding! up to 4x decode …

X AI KOLs Timeline ↗ · 2026-08-20 Cached

Liquid AI releases DSpark draft models for their LFM series, incorporating speculative decoding to achieve up to 4x decode speedup on device while maintaining output quality.

0 favorites 0 likes
#speculative-decoding

Up to 3.2x Faster Inference with LFM2.5-DSpark

Hugging Face Blog ↗ · 2026-08-20 Cached

Liquid AI releases DSpark draft model checkpoints for the LFM2.5 family, enabling up to 3.2x faster inference on GPUs and devices with minimal quality trade-off, and with day-one support for open-source tools like llama.cpp and SGLang.

0 favorites 0 likes
#speculative-decoding

3 days benchmarking most llama.cpp flags on my weird 40gb vram laptop + tb4 egpu setup. Got +70% generation, +40% prefill, 60k more context, and filed a bug in llama around MTP. What I learned.

Reddit r/LocalLLaMA ↗ · 2026-08-20

The author benchmarked llama.cpp flags on a hybrid GPU setup with an RTX 4090 laptop and AMD XTX 7900 eGPU, achieving 70% faster generation, 40% faster prefill, and discovering a bug related to MTP in multi-GPU configurations.

0 favorites 0 likes
#speculative-decoding

Accelerating Visual On-Policy Distillation with Batched Speculative Jacobi Rollouts

arXiv cs.LG ↗ · 2026-08-20 Cached

This paper introduces HB-SJD, a batched speculative Jacobi decoding method for visual on-policy distillation that accelerates rollout generation by processing multiple tokens in parallel, reducing training time while preserving generation quality.

0 favorites 0 likes
#speculative-decoding

NVFP4 on VOLTA! Despite being built for Blackwell, I made four 2017 V100s run Qwen 3.8 NVFP4 natively and match my $6000 RTX 5090.

Reddit r/LocalLLaMA ↗ · 2026-08-19

The author achieved performance parity between four 2017 Tesla V100 GPUs and a modern RTX 5090 when running the Qwen 3.8 model with NVFP4 precision, using custom software optimizations like the QPN kernel for efficient inference.

0 favorites 0 likes
#speculative-decoding

DFlash 2: Keep Drafting Parallel

Reddit r/LocalLLaMA ↗ · 2026-08-18 Cached

DFlash 2 improves speculative decoding by predicting tokens in parallel, achieving over 20% more output per verification pass with minimal latency, and is integrated into major inference engines like SGLang and vLLM.

0 favorites 0 likes
#speculative-decoding

DeepSeek V4 Flash 0731 on Strix Halo: draft model, n_max sweep, and a launch line that actually helps

Reddit r/LocalLLaMA ↗ · 2026-08-18

The article presents benchmark results for DeepSeek V4 Flash 0731 on Strix Halo hardware, showing performance with different draft models and n_max settings, concluding that n_max=3 offers the best speed balance.

0 favorites 0 likes
#speculative-decoding

incoai/Qwen3.8-27B-DFlash2

Hugging Face Models Trending ↗ · 2026-08-18 Cached

DFlash 2 is a block-diffusion draft model for speculative decoding that improves inference speed for the Qwen3.8-27B language model, offering higher acceptance length and throughput in benchmarks.

0 favorites 0 likes
#speculative-decoding

Qwen3.8-27B at 256K on a 24GB RTX PRO 4000 SFF (432 GB/s): 50 tok/s with MTP

Hacker News Top ↗ · 2026-08-17 Cached

The article details an experiment achieving 50 tokens per second inference with Qwen3.8-27B at 256K context on a 24GB GPU using Multi-Token Prediction and custom optimizations.

0 favorites 0 likes
#speculative-decoding

Qwen3.8-27B Q8_0 on Strix Halo is seriously impressive

Reddit r/LocalLLaMA ↗ · 2026-08-17

The author tested the Qwen3.8-27B Q8_0 AI model on a ROG Flow Z13 with Ryzen AI Max+ 395, achieving impressive local inference performance in generating a flight simulator using Lemonade Server and llama.cpp with speculative decoding.

0 favorites 0 likes
#speculative-decoding

Qwen3.8-27B-int4-AutoRound (18GB) - with working MTP spec decode

Reddit r/LocalLLaMA ↗ · 2026-08-16 Cached

This article presents a quantized version of the Qwen3.8-27B model using INT4 AutoRound quantization with working MTP for speculative decoding, achieving significant inference speedups on GPUs.

0 favorites 0 likes
#speculative-decoding

z-lab/Qwen3.8-27B-DFlash2

Hugging Face Models Trending ↗ · 2026-08-15 Cached

Introduces DFlash 2, a block-diffusion drafter for speculative decoding with the Qwen3.8-27B model, demonstrating improved acceptance length and throughput in benchmarks.

0 favorites 0 likes
#speculative-decoding

Qwen3.8-27B is now up to ~3× faster on Apple Silicon with mlx-dspark

Reddit r/LocalLLaMA ↗ · 2026-08-14

mlx-dspark v0.10.0 adds support for Qwen3.8-27B on Apple Silicon, providing up to 3x faster inference through speculative decoding with lossless verification.

0 favorites 0 likes
#speculative-decoding

NInfer day0 support for Qwen3.8 27b: ~200 tok/s generation, with tons of engine improvments

Reddit r/LocalLLaMA ↗ · 2026-08-14

NInfer adds day-0 support for the Qwen3.8-27B model, achieving around 200 tokens per second on a single RTX 5090 with speculative decoding, and includes engine improvements like concurrent requests and kernel optimizations.

0 favorites 0 likes
#speculative-decoding

FlashDrive: Flash Vision-Language-Action Inference for Autonomous Driving

arXiv cs.AI ↗ · 2026-08-14 Cached

FlashDrive is an algorithm-system co-design framework that cuts the inference latency of vision-language-action models for autonomous driving by 4.7× (from 717 ms to 151 ms on a single GPU) using streaming KV-cache reuse, non-autoregressive diffusion drafting, and adaptive step caching, with negligible accuracy loss.

0 favorites 0 likes
#speculative-decoding

Decoupled Contrastive Decoding via Expert-Aligned Drafting

arXiv cs.CL ↗ · 2026-08-14 Cached

This paper introduces Decoupled Contrastive Decoding (DCD), which uses an expert-aligned lightweight proposer for speculative decoding while keeping the contrastive signal only in verification, achieving speedups over vanilla contrastive decoding without degrading output distribution.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback