Tag
A talk about llama.cpp speculative decoding methods (MTP, dflash, dspark) at the dotAI conference will have its replay available soon.
Intel releases OpenVINO 2026.4 with expanded AI model support, performance enhancements like multi-token prediction, and new features for profiling and inference across CPUs, GPUs, and NPUs.
This paper constructs a cost-quality-latency Pareto atlas for LLM inference optimizations, using a calibrated simulator to evaluate configurations and combinations across different hardware and regimes.
FlexEE introduces a self-speculative and KV-cache-compatible early exiting framework for efficient LLM inference in offloading deployments, achieving significant speedups on Llama models with minimal accuracy degradation.
The author optimized DeepSeek V4.1 Flash for Apple M3 Ultra, achieving up to 40 t/s decode speed with DSpark speculative decoding while maintaining byte-identical accuracy to the upstream model.
R9V update delivers approximately 100 tokens per second on Qwen3.8 Flash Next with IQ4_XS on dual AMD R9700 GPUs, adds support for Q4_K_XL with 50 tok/s, fixes crashes, and enhances diagnostics.
This paper investigates the losslessness of Orthrus, a hybrid autoregressive-diffusion model for inference acceleration, finding that it requires high numerical precision (FP32) for exact trajectory matching, while BF16 divergence does not impair downstream performance.
A draft model for Qwen 3.8 27B is highly effective on 16 GB GPUs, achieving around 60 tokens per second on an RX 9070 XT and offering better VRAM efficiency than built-in MTP.
The Qwen 3.8 Flash Next model achieved a 100% score on the CUDA.fast benchmark with a 117.8% composite speed increase on DGX Spark, utilizing speculative decoding.
Osprey introduces a target-agnostic pre-training method for drafters in speculative decoding, improving efficiency by bootstrapping from off-the-shelf models and adapting with minimal target-specific work, achieving significant acceptance rate improvements across multiple LLMs.
This paper introduces X-CoSD, a communication-efficient cross-vocabulary collaborative speculative decoding framework that optimizes distributed LLM inference by splitting residual resampling to reduce overhead while preserving server LLM quality.
The author explains how speculative decoding affects coding agent speed, with higher acceptance rates on boilerplate code leading to faster typing, and discusses other factors like cache misses that impact performance.
Benchmarks comparing SGLang, llama.cpp, and FreeToken on Qwen3.8-Flash-Next at full context show SGLang achieves the fastest time to first token at 35.4s, while llama.cpp baseline takes 258.4s, with speculative decoding providing performance improvements.
Benchmark tests on Ling-3.0-flash show that higher acceptance length in multi-token prediction decreases prose throughput, with n=1 being the most efficient setting for the evaluated workloads.
This paper presents a system to accelerate speculative decoding in large-scale RL post-training through online draft co-training, using context-parallel attention extensions and cross-stage feature transport.
Jina-OCR-v1 is an efficient end-to-end document parsing model that uses speculative decoding and dense verifiable rewards to achieve high accuracy and speed on low-budget GPUs, scoring 91.14 on OmniDocBench v1.6 and 83.4 on olmOCR-Bench.
AdaptiveSpec is a training-free per-step speculative decoding method that adaptively adjusts token verification and draft tree shape to enhance LLM inference throughput, improving performance by up to 56% while maintaining high accuracy across benchmarks.
TorchSpec is a PyTorch-native framework for training speculative decoding draft models, released in collaboration with vllm and demonstrated with Kimi K3 draft models on NVIDIA GB200 hardware.
Release of a speculative decoding implementation in Uzu, initially supporting Qwen3.6 27B with upcoming support for Qwen3.8 27B and Muse Glimmer.
User @ViC305 successfully runs DeepSeek-V4-Flash-Vision with EXL3 MixedK and DSpark speculative decoding on a single DGX Spark, achieving improved performance and fixing technical issues for multimodal AI deployment.