Tag
MegaCapybara 是一个针对 NVIDIA RTX 5090 优化的开源 LLM 推理引擎,支持 Qwen3.8-27B,单智能体可达 500+ tokens/s,多智能体并发最高 2000+ tokens/s,具备投机解码、共享上下文池和智能 VRAM/RAM/磁盘缓存管理。
作者在单张 RTX PRO 6000 上对三款 Qwen3.8 27B 检查点进行 10 项统一测试,Flash-Next 赢下多数任务,而 DFlash2 投机解码将 27B 的推理速度从 75 提升到 210 tok/s(2.8 倍),并对比了 SGLang 与 vLLM 在不同量化导出下的性能差异。
A quantized and abliterated Qwen3.8-27B CODER model variant (IQ4_XS with imatrix) released in a GGUF build optimized by LexiPanel to fit a single 24 GB GPU with a 262k context window, featuring an MTP head for speculative decoding at 85% draft acceptance on an RX 7900 XTX.
作者对开源 LLM 推理服务 Gufo 0.4.0 进行实测复现:其宣传的 70 tok/s 确实存在,但主要来自重复词 prompt 下的投机解码命中率,普通 prompt 上约 39 tok/s;与 halogen 相比,日常生成速度慢约 13-18%,但长 prompt 处理更快。
作者发布 Slipstream——一款针对 Apple Silicon 的 C++ Metal 推理引擎,支持原生 SSD 专家流式加载与投机草稿解码,可在 64GB Mac 上以 41–52 tok/s 运行 95.5 GiB 的 Qwen3.8-Flash-Next 模型,比 llama.cpp 快 1.76 倍,且 130k 长上下文下速度不衰减。
This paper proposes a rollout-based training framework (Anchor-Label Relabelling and In-Rollout Anchors) to recover full supervision for speculative decoding block drafters trained on off-policy corpora, boosting greedy accepted length by up to 36.5% over DFlash without modifying the training text.
DEdit introduces a diffusion-based drafter for speculative decoding that iteratively edits drafts via token-to-token predictions, using a ProposalMix training scheme to repair errors while preserving correct tokens. It achieves macro-average speedups of 5.72× and 5.97× on Qwen3-4B and Qwen3-8B across seven benchmarks.
AgSpec 是一个检索式推测解码框架,通过从会话、工作区和全局语料中检索草稿并动态调整草稿长度,在多智能体编码基准上实现了最高 4.37 倍(batch size 1)和 4.76 倍(batch size 16)的生成吞吐提升,优于五种现有检索式草稿器和 EAGLE-3。
A same-harness benchmark on a Mac Studio M5 Ultra 256GB compares Qwen3.8-Flash-Next (182GB oQ8e) vs Laguna-S-2.1 GGUF at up to 262K context. Qwen sustains ~4,200 tok/s linear prefill and stable decode, while Laguna degrades superlinearly; quality is a 4/4 draw with very different answer styles.
A head-to-head benchmark of two community inference stacks running the 320B GLM-5.3-Flash on a pair of NVIDIA DGX Sparks, comparing Entrpi's EXL3 setup against TensorFold with DFlash2 drafters across prose, code, and prefill workloads.
A user reports that the open-source TensorFold inference engine, running a 4-bit Qwen3.8-27B with a DFlash2 drafting model, reaches 40-60 tok/s on an M5 Pro Mac mini, beating MTPLX and rivaling an RTX 3090 Ti. This highlights rapid performance gains for local LLM inference on Apple Silicon.
Artificial Analysis open-sources AA-AgentPerf-Local, an inference benchmarking tool that replays real agent trajectories to measure how fast agentic AI runs on laptops and workstations, with initial results for NVIDIA DGX Spark, RTX 5090, AMD Ryzen AI Halo, and MacBook Pro M5 Pro.
提出DSpine,一种通过网络深度注入相邻因果条件的并行推测解码drafter,在SGLang中实现高效并行执行,较DFlash在Qwen3-8B上平均接受长度提升27.8%、吞吐量提升23.3%。
ninfer-ext is an extended fork of NInfer that adds support for larger Qwen models, faster speculative decoding, agent serving features, and EXL3 quantization for Qwen3.8-27B on NVIDIA RTX 5090.
Ornith-1.5 models integrated with DFlash draft models for speculative decoding have been released on Hugging Face in 9B, 397B, and 35B-A3B sizes.
This paper introduces mentored decoding, a formal approach to lossy speculative decoding that improves inference speed in language models by allowing controlled divergence from the target model, connecting it to boosting theory and proving key properties for optimization.
A blog post detailing performance optimizations in llama.cpp that make prompt lookup drafting up to 42x faster and reduce memory usage by 2.6x, based on techniques from Daniel Lemire and Martin Ankerl.
WaveFront Decoding introduces a training-free self-speculative decoding framework for looped language models that reduces latency by concurrently batching drafting and verification, achieving up to 4.81x speedup on Huginn-3.5B.
BeeLlama v0.4.7 is released, offering gigabytes of VRAM savings through independent drafter ubatch sizing for MTP and DFlash, along with major ROCm and CUDA performance enhancements for AI inference.
Today, we release an experimental DSpark draft model for our vision-language model (VLM) LFM2.5-VL-3B, which adds a speculative decoding path for faster inference with minimal memory cost.