speculative-decoding

Tag

Cards List
#speculative-decoding

The fastest interference engine for RTX5090 and Qwen3.8 27B. Twice as fast as ninfer. 500+ t/s single coding, 2000+t/s up to 12 agents at the same time with 800k context. Smart VRAM-RAM-DISC Cache management, Loop Guard, Nice UI etc.

Reddit r/LocalLLaMA ↗ · 3h ago Cached

MegaCapybara 是一个针对 NVIDIA RTX 5090 优化的开源 LLM 推理引擎,支持 Qwen3.8-27B,单智能体可达 500+ tokens/s,多智能体并发最高 2000+ tokens/s,具备投机解码、共享上下文池和智能 VRAM/RAM/磁盘缓存管理。

0 favorites 0 likes
#speculative-decoding

I spent 3 weeks testing local Qwen3.8 on the new low-latency SGLang/vLLM recipes: DFlash2 2.8x. Builds: RadixArk + Inferact 27B NVFP4, 27B BF16, orcarouter 27B Uncensored, Flash-Next NVFP4

Reddit r/LocalLLaMA ↗ · 7h ago

作者在单张 RTX PRO 6000 上对三款 Qwen3.8 27B 检查点进行 10 项统一测试,Flash-Next 赢下多数任务,而 DFlash2 投机解码将 27B 的推理速度从 75 提升到 210 tok/s(2.8 倍),并对比了 SGLang 与 vLLM 在不同量化导出下的性能差异。

0 favorites 0 likes
#speculative-decoding

Tuned/abliterated Qwen3.8-27b into a 24gb card 262k guff using the newest unreleased version of LexiPanel. It's fast with reliable draft acceptance. Made for 7900xtx but should work on whatever 24gb card with this setup and headless. Doesn't get dumber while coding like most of the other fine-tunes.

Reddit r/LocalLLaMA ↗ · 15h ago

A quantized and abliterated Qwen3.8-27B CODER model variant (IQ4_XS with imatrix) released in a GGUF build optimized by LexiPanel to fit a single 24 GB GPU with a 262k context window, featuring an MTP head for speculative decoding at 85% draft acceptance on an RX 7900 XTX.

0 favorites 0 likes
#speculative-decoding

Gufo performance .... 70tps Qwen 3.8 27b but you need to read the fine print.

Reddit r/LocalLLaMA ↗ · yesterday

作者对开源 LLM 推理服务 Gufo 0.4.0 进行实测复现:其宣传的 70 tok/s 确实存在,但主要来自重复词 prompt 下的投机解码命中率,普通 prompt 上约 39 tok/s;与 halogen 相比,日常生成速度慢约 13-18%,但长 prompt 处理更快。

0 favorites 0 likes
#speculative-decoding

Running 95.5 GiB Qwen3.8-Flash-Next at 41–52 tok/s on a 64GB Mac (1.76x faster than llama.cpp): Slipstream release, 130k context scaling, + Swift variant

Reddit r/LocalLLaMA ↗ · yesterday

作者发布 Slipstream——一款针对 Apple Silicon 的 C++ Metal 推理引擎,支持原生 SSD 专家流式加载与投机草稿解码,可在 64GB Mac 上以 41–52 tok/s 运行 95.5 GiB 的 Qwen3.8-Flash-Next 模型,比 llama.cpp 快 1.76 倍,且 130k 长上下文下速度不衰减。

0 favorites 0 likes
#speculative-decoding

Recovering Off-Policy Supervision for Speculative Decoding

arXiv cs.CL ↗ · 2d ago Cached

This paper proposes a rollout-based training framework (Anchor-Label Relabelling and In-Rollout Anchors) to recover full supervision for speculative decoding block drafters trained on off-policy corpora, boosting greedy accepted length by up to 36.5% over DFlash without modifying the training text.

0 favorites 0 likes
#speculative-decoding

DEdit: Iterative Draft Editing for Speculative Decoding

arXiv cs.CL ↗ · 2d ago Cached

DEdit introduces a diffusion-based drafter for speculative decoding that iteratively edits drafts via token-to-token predictions, using a ProposalMix training scheme to repair errors while preserving correct tokens. It achieves macro-average speedups of 5.72× and 5.97× on Qwen3-4B and Qwen3-8B across seven benchmarks.

0 favorites 0 likes
#speculative-decoding

AgSpec: Pushing the Limits of Retrieval-Based Speculative Decoding in Coding Agent Pipelines

Hugging Face Daily Papers ↗ · 2d ago Cached

AgSpec 是一个检索式推测解码框架,通过从会话、工作区和全局语料中检索草稿并动态调整草稿长度,在多智能体编码基准上实现了最高 4.37 倍(batch size 1)和 4.76 倍(batch size 16)的生成吞吐提升,优于五种现有检索式草稿器和 EAGLE-3。

0 favorites 0 likes
#speculative-decoding

M5 Ultra - Qwen3.8 Flash Next vs Laguna S 2.1

Reddit r/LocalLLaMA ↗ · 2d ago

A same-harness benchmark on a Mac Studio M5 Ultra 256GB compares Qwen3.8-Flash-Next (182GB oQ8e) vs Laguna-S-2.1 GGUF at up to 262K context. Qwen sustains ~4,200 tok/s linear prefill and stable decode, while Laguna degrades superlinearly; quality is a 4/4 draw with very different answer styles.

0 favorites 0 likes
#speculative-decoding

@no_stp_on_snek: Same two GB10s. GLM-5.3-Flash, two stacks. One streamed call each, temperature 0.2, 384 tokens, prompts around 3.7k tok…

X AI KOLs Timeline ↗ · 2d ago Cached

A head-to-head benchmark of two community inference stacks running the 320B GLM-5.3-Flash on a pair of NVIDIA DGX Sparks, comparing Entrpi's EXL3 setup against TensorFold with DFlash2 drafters across prose, code, and prefill workloads.

0 favorites 0 likes
#speculative-decoding

Tensorfold runs Qwen3.8-27B really well on m5 pro mac mini, tps beats MTPLX

Reddit r/LocalLLaMA ↗ · 2d ago

A user reports that the open-source TensorFold inference engine, running a 4-bit Qwen3.8-27B with a DFlash2 drafting model, reaches 40-60 tok/s on an M5 Pro Mac mini, beating MTPLX and rivaling an RTX 3090 Ti. This highlights rapid performance gains for local LLM inference on Apple Silicon.

0 favorites 0 likes
#speculative-decoding

AA-AgentPerf-Local: Benchmarking local AI agents on laptops and workstations

Reddit r/LocalLLaMA ↗ · 2d ago Cached

Artificial Analysis open-sources AA-AgentPerf-Local, an inference benchmarking tool that replays real agent trajectories to measure how fast agentic AI runs on laptops and workstations, with initial results for NVIDIA DGX Spark, RTX 5090, AMD Ryzen AI Halo, and MacBook Pro M5 Pro.

0 favorites 0 likes
#speculative-decoding

Draft in Parallel, Condition Through Depth: Adjacent Causal Injection for Speculative Decoding

arXiv cs.LG ↗ · 3d ago Cached

提出DSpine,一种通过网络深度注入相邻因果条件的并行推测解码drafter,在SGLang中实现高效并行执行,较DFlash在Qwen3-8B上平均接受长度提升27.8%、吞吐量提升23.3%。

0 favorites 0 likes
#speculative-decoding

exl3 now in ninfer-ext

Reddit r/LocalLLaMA ↗ · 3d ago Cached

ninfer-ext is an extended fork of NInfer that adds support for larger Qwen models, faster speculative decoding, agent serving features, and EXL3 quantization for Qwen3.8-27B on NVIDIA RTX 5090.

0 favorites 0 likes
#speculative-decoding

Ornith-1.5 DFlash

Reddit r/LocalLLaMA ↗ · 3d ago

Ornith-1.5 models integrated with DFlash draft models for speculative decoding have been released on Hugging Face in 9B, 397B, and 35B-A3B sizes.

0 favorites 0 likes
#speculative-decoding

Mentored Decoding: Faster Inference meets Boosting

arXiv cs.LG ↗ · 4d ago Cached

This paper introduces mentored decoding, a formal approach to lossy speculative decoding that improves inference speed in language models by allowing controlled divergence from the target model, connecting it to boosting theory and proving key properties for optimization.

0 favorites 0 likes
#speculative-decoding

42x Faster Prompt Lookup Drafting in llama.cpp

Reddit r/LocalLLaMA ↗ · 6d ago Cached

A blog post detailing performance optimizations in llama.cpp that make prompt lookup drafting up to 42x faster and reduce memory usage by 2.6x, based on techniques from Daniel Lemire and Martin Ankerl.

0 favorites 0 likes
#speculative-decoding

WaveFront Decoding: Parallelized Self-Speculative Decoding for Looped Language Models

Hugging Face Daily Papers ↗ · 2026-09-26 Cached

WaveFront Decoding introduces a training-free self-speculative decoding framework for looped language models that reduces latency by concurrently batching drafting and verification, achieving up to 4.81x speedup on Huginn-3.5B.

0 favorites 0 likes
#speculative-decoding

@Anbeeld: BeeLlama v0.4.7 is out! You can now save gigabytes of VRAM when using MTP and DFlash! In the example shown in the image…

X AI KOLs Timeline ↗ · 2026-09-25

BeeLlama v0.4.7 is released, offering gigabytes of VRAM savings through independent drafter ubatch sizing for MTP and DFlash, along with major ROCm and CUDA performance enhancements for AI inference.

0 favorites 0 likes
#speculative-decoding

Accelerating vision-language models with LFM2.5-VL-DSpark

Hugging Face Blog ↗ · 2026-09-24 Cached

Today, we release an experimental DSpark draft model for our vision-language model (VLM) LFM2.5-VL-3B, which adds a speculative decoding path for faster inference with minimal memory cost.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback