Tag
A user reports that the open-source TensorFold inference engine, running a 4-bit Qwen3.8-27B with a DFlash2 drafting model, reaches 40-60 tok/s on an M5 Pro Mac mini, beating MTPLX and rivaling an RTX 3090 Ti. This highlights rapid performance gains for local LLM inference on Apple Silicon.
Artificial Analysis open-sources AA-AgentPerf-Local, an inference benchmarking tool that replays real agent trajectories to measure how fast agentic AI runs on laptops and workstations, with initial results for NVIDIA DGX Spark, RTX 5090, AMD Ryzen AI Halo, and MacBook Pro M5 Pro.
提出DSpine,一种通过网络深度注入相邻因果条件的并行推测解码drafter,在SGLang中实现高效并行执行,较DFlash在Qwen3-8B上平均接受长度提升27.8%、吞吐量提升23.3%。
ninfer-ext is an extended fork of NInfer that adds support for larger Qwen models, faster speculative decoding, agent serving features, and EXL3 quantization for Qwen3.8-27B on NVIDIA RTX 5090.
Ornith-1.5 models integrated with DFlash draft models for speculative decoding have been released on Hugging Face in 9B, 397B, and 35B-A3B sizes.
This paper introduces mentored decoding, a formal approach to lossy speculative decoding that improves inference speed in language models by allowing controlled divergence from the target model, connecting it to boosting theory and proving key properties for optimization.
A blog post detailing performance optimizations in llama.cpp that make prompt lookup drafting up to 42x faster and reduce memory usage by 2.6x, based on techniques from Daniel Lemire and Martin Ankerl.
WaveFront Decoding introduces a training-free self-speculative decoding framework for looped language models that reduces latency by concurrently batching drafting and verification, achieving up to 4.81x speedup on Huginn-3.5B.
BeeLlama v0.4.7 is released, offering gigabytes of VRAM savings through independent drafter ubatch sizing for MTP and DFlash, along with major ROCm and CUDA performance enhancements for AI inference.
Today, we release an experimental DSpark draft model for our vision-language model (VLM) LFM2.5-VL-3B, which adds a speculative decoding path for faster inference with minimal memory cost.
DPara is a parallel speculative decoding framework that eliminates probabilistic fallback by precomputing draft representations, achieving average speedups of 3.21× to 3.52× over autoregressive decoding on Qwen3 models.
The tweet discusses how providers tune quantization and speculative decoding strategies post-launch based on usage patterns and hardware, and introduces an interactive tutorial on speculative decoding for the NeurIPS Education Track.
The paper proposes TSS, a target-side sparsification framework for speculative decoding in domain-specific LLMs, which improves inference efficiency and performance by skipping selected target layers.
A developer shares a fork of FreeToken, an edge-native MoE serving engine, with added support for DeepSeek-V4.1, vision capabilities for Qwen models, and speculative decoding, including benchmarks on dual RTX 3090 hardware.
This paper formulates speculative decoding in Mixture-of-Experts models as a Stochastic Shortest Path problem and uses a diagnostic Oracle to demonstrate that optimal decisions follow a necessary condition balancing cost and progress.
TreeSpark introduces calibrated, load-adaptive draft trees for semi-autoregressive speculative decoding in language models, increasing draft token acceptance by 15-25% and speeding up decoding by 8-14%.
Extended Splash Engine to support native 8-bit Qwen3.8-27B on Apple Silicon, achieving 37-55 tok/s without quantization degradation and scaling up to 256k context.
SwitchSD is an adaptive framework for speculative decoding in LLMs that uses intrinsic model signals to dynamically switch between neural drafting and context-based copying, achieving up to 15% throughput gains over baselines like EAGLE3.
Zarya is a hybrid language model that jointly optimizes autoregressive and masked diffusion objectives for flexible training and dual-mode inference, with publicly released models in sizes 0.6B, 1.7B, and 4B.
ByteShape releases full ShapeLearn quantized versions of the Qwen 3.8 27B model in GGUF format, with benchmarking showing improvements in quality-speed frontier and support for speculative decoding.