speculative-decoding

Tag

Cards List
#speculative-decoding

LFM2.5 2.6b vs MiniCPM5 2b

Reddit r/LocalLLaMA ↗ · 4h ago

作者在 M1 Air 上使用 llama-server 对比了 LFM2.5 2.6B 与 MiniCPM5 2B 在小型 agentic 工具调用任务上的表现。结论是 LFM2.5 明显更优——尽管参数更多但速度更快、内存占用更低,而 MiniCPM5 虽然在 agentic 任务上可能更强,但常常在英文提问下返回中文。

0 favorites 0 likes
#speculative-decoding

@servasyy_ai: https://x.com/servasyy_ai/status/2106563331976991217

X AI KOLs Timeline ↗ · 20h ago Cached

This long post is a real-world test of running Qwen3.8-27B on two second-hand V100 32G GPUs. By rewriting the kernel of the inference engine NInfer, adding dual-GPU support and speculative decoding, generation speed on a 200K-token code task jumped from 36 tok/s to roughly 100 tok/s, beating the 80.9 tok/s of a single 4090 48G. The post also analyzes how VRAM bandwidth, quantization, and NVLink affect inference performance.

0 favorites 0 likes
#speculative-decoding

I built Ninfer 4080 for 16GB class GPUs

Reddit r/LocalLLaMA ↗ · yesterday

A community developer released Ninfer 4080, an optimized local inference runtime that runs the Qwen 3.8 27B GSQ model at 100k context on a 16GB RTX 4080, achieving up to ~2720 tok/s prefill and 262 tok/s decode via DFlash2 speculative decoding — significantly faster than general-purpose engines like llama.cpp or vLLM.

0 favorites 0 likes
#speculative-decoding

The fastest interference engine for RTX5090 and Qwen3.8 27B. Twice as fast as ninfer. 500+ t/s single coding, 2000+t/s up to 12 agents at the same time with 800k context. Smart VRAM-RAM-DISC Cache management, Loop Guard, Nice UI etc.

Reddit r/LocalLLaMA ↗ · yesterday Cached

MegaCapybara 是一个针对 NVIDIA RTX 5090 优化的开源 LLM 推理引擎,支持 Qwen3.8-27B,单智能体可达 500+ tokens/s,多智能体并发最高 2000+ tokens/s,具备投机解码、共享上下文池和智能 VRAM/RAM/磁盘缓存管理。

0 favorites 0 likes
#speculative-decoding

@volatilemarkts: Same Mac. Same model. Same words. My 512 GB M3 Ultra ran GLM-5.3-Flash on MLX last week: 27 tok/s, first token in 0.46 …

X AI KOLs Timeline ↗ · yesterday Cached

TensorFold 0.6.2 is an Apache 2.0 open-source LLM inference engine claiming over 2x faster token throughput than MLX on Apple Silicon (61 vs 27 tok/s on an M3 Ultra) via multi-token-prediction speculative decoding, with parallel concurrent requests and OpenAI-compatible API.

0 favorites 0 likes
#speculative-decoding

I spent 3 weeks testing local Qwen3.8 on the new low-latency SGLang/vLLM recipes: DFlash2 2.8x. Builds: RadixArk + Inferact 27B NVFP4, 27B BF16, orcarouter 27B Uncensored, Flash-Next NVFP4

Reddit r/LocalLLaMA ↗ · 2d ago

作者在单张 RTX PRO 6000 上对三款 Qwen3.8 27B 检查点进行 10 项统一测试,Flash-Next 赢下多数任务,而 DFlash2 投机解码将 27B 的推理速度从 75 提升到 210 tok/s(2.8 倍),并对比了 SGLang 与 vLLM 在不同量化导出下的性能差异。

0 favorites 0 likes
#speculative-decoding

Tuned/abliterated Qwen3.8-27b into a 24gb card 262k guff using the newest unreleased version of LexiPanel. It's fast with reliable draft acceptance. Made for 7900xtx but should work on whatever 24gb card with this setup and headless. Doesn't get dumber while coding like most of the other fine-tunes.

Reddit r/LocalLLaMA ↗ · 2d ago

A quantized and abliterated Qwen3.8-27B CODER model variant (IQ4_XS with imatrix) released in a GGUF build optimized by LexiPanel to fit a single 24 GB GPU with a 262k context window, featuring an MTP head for speculative decoding at 85% draft acceptance on an RX 7900 XTX.

0 favorites 0 likes
#speculative-decoding

Gufo performance .... 70tps Qwen 3.8 27b but you need to read the fine print.

Reddit r/LocalLLaMA ↗ · 3d ago

作者对开源 LLM 推理服务 Gufo 0.4.0 进行实测复现:其宣传的 70 tok/s 确实存在,但主要来自重复词 prompt 下的投机解码命中率,普通 prompt 上约 39 tok/s;与 halogen 相比,日常生成速度慢约 13-18%,但长 prompt 处理更快。

0 favorites 0 likes
#speculative-decoding

Running 95.5 GiB Qwen3.8-Flash-Next at 41–52 tok/s on a 64GB Mac (1.76x faster than llama.cpp): Slipstream release, 130k context scaling, + Swift variant

Reddit r/LocalLLaMA ↗ · 3d ago

作者发布 Slipstream——一款针对 Apple Silicon 的 C++ Metal 推理引擎,支持原生 SSD 专家流式加载与投机草稿解码,可在 64GB Mac 上以 41–52 tok/s 运行 95.5 GiB 的 Qwen3.8-Flash-Next 模型,比 llama.cpp 快 1.76 倍,且 130k 长上下文下速度不衰减。

0 favorites 0 likes
#speculative-decoding

Recovering Off-Policy Supervision for Speculative Decoding

arXiv cs.CL ↗ · 3d ago Cached

This paper proposes a rollout-based training framework (Anchor-Label Relabelling and In-Rollout Anchors) to recover full supervision for speculative decoding block drafters trained on off-policy corpora, boosting greedy accepted length by up to 36.5% over DFlash without modifying the training text.

0 favorites 0 likes
#speculative-decoding

DEdit: Iterative Draft Editing for Speculative Decoding

arXiv cs.CL ↗ · 3d ago Cached

DEdit introduces a diffusion-based drafter for speculative decoding that iteratively edits drafts via token-to-token predictions, using a ProposalMix training scheme to repair errors while preserving correct tokens. It achieves macro-average speedups of 5.72× and 5.97× on Qwen3-4B and Qwen3-8B across seven benchmarks.

0 favorites 0 likes
#speculative-decoding

AgSpec: Pushing the Limits of Retrieval-Based Speculative Decoding in Coding Agent Pipelines

Hugging Face Daily Papers ↗ · 3d ago Cached

AgSpec 是一个检索式推测解码框架,通过从会话、工作区和全局语料中检索草稿并动态调整草稿长度,在多智能体编码基准上实现了最高 4.37 倍(batch size 1)和 4.76 倍(batch size 16)的生成吞吐提升,优于五种现有检索式草稿器和 EAGLE-3。

0 favorites 0 likes
#speculative-decoding

M5 Ultra - Qwen3.8 Flash Next vs Laguna S 2.1

Reddit r/LocalLLaMA ↗ · 4d ago

A same-harness benchmark on a Mac Studio M5 Ultra 256GB compares Qwen3.8-Flash-Next (182GB oQ8e) vs Laguna-S-2.1 GGUF at up to 262K context. Qwen sustains ~4,200 tok/s linear prefill and stable decode, while Laguna degrades superlinearly; quality is a 4/4 draw with very different answer styles.

0 favorites 0 likes
#speculative-decoding

@no_stp_on_snek: Same two GB10s. GLM-5.3-Flash, two stacks. One streamed call each, temperature 0.2, 384 tokens, prompts around 3.7k tok…

X AI KOLs Timeline ↗ · 4d ago Cached

A head-to-head benchmark of two community inference stacks running the 320B GLM-5.3-Flash on a pair of NVIDIA DGX Sparks, comparing Entrpi's EXL3 setup against TensorFold with DFlash2 drafters across prose, code, and prefill workloads.

0 favorites 0 likes
#speculative-decoding

Tensorfold runs Qwen3.8-27B really well on m5 pro mac mini, tps beats MTPLX

Reddit r/LocalLLaMA ↗ · 4d ago

A user reports that the open-source TensorFold inference engine, running a 4-bit Qwen3.8-27B with a DFlash2 drafting model, reaches 40-60 tok/s on an M5 Pro Mac mini, beating MTPLX and rivaling an RTX 3090 Ti. This highlights rapid performance gains for local LLM inference on Apple Silicon.

0 favorites 0 likes
#speculative-decoding

AA-AgentPerf-Local: Benchmarking local AI agents on laptops and workstations

Reddit r/LocalLLaMA ↗ · 4d ago Cached

Artificial Analysis open-sources AA-AgentPerf-Local, an inference benchmarking tool that replays real agent trajectories to measure how fast agentic AI runs on laptops and workstations, with initial results for NVIDIA DGX Spark, RTX 5090, AMD Ryzen AI Halo, and MacBook Pro M5 Pro.

0 favorites 0 likes
#speculative-decoding

Draft in Parallel, Condition Through Depth: Adjacent Causal Injection for Speculative Decoding

arXiv cs.LG ↗ · 4d ago Cached

提出DSpine,一种通过网络深度注入相邻因果条件的并行推测解码drafter,在SGLang中实现高效并行执行,较DFlash在Qwen3-8B上平均接受长度提升27.8%、吞吐量提升23.3%。

0 favorites 0 likes
#speculative-decoding

exl3 now in ninfer-ext

Reddit r/LocalLLaMA ↗ · 4d ago Cached

ninfer-ext is an extended fork of NInfer that adds support for larger Qwen models, faster speculative decoding, agent serving features, and EXL3 quantization for Qwen3.8-27B on NVIDIA RTX 5090.

0 favorites 0 likes
#speculative-decoding

Ornith-1.5 DFlash

Reddit r/LocalLLaMA ↗ · 5d ago

Ornith-1.5 models integrated with DFlash draft models for speculative decoding have been released on Hugging Face in 9B, 397B, and 35B-A3B sizes.

0 favorites 0 likes
#speculative-decoding

Mentored Decoding: Faster Inference meets Boosting

arXiv cs.LG ↗ · 5d ago Cached

This paper introduces mentored decoding, a formal approach to lossy speculative decoding that improves inference speed in language models by allowing controlled divergence from the target model, connecting it to boosting theory and proving key properties for optimization.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback