Tag
The paper proposes TSS, a target-side sparsification framework for speculative decoding in domain-specific LLMs, which improves inference efficiency and performance by skipping selected target layers.
A developer shares a fork of FreeToken, an edge-native MoE serving engine, with added support for DeepSeek-V4.1, vision capabilities for Qwen models, and speculative decoding, including benchmarks on dual RTX 3090 hardware.
This paper formulates speculative decoding in Mixture-of-Experts models as a Stochastic Shortest Path problem and uses a diagnostic Oracle to demonstrate that optimal decisions follow a necessary condition balancing cost and progress.
TreeSpark introduces calibrated, load-adaptive draft trees for semi-autoregressive speculative decoding in language models, increasing draft token acceptance by 15-25% and speeding up decoding by 8-14%.
Extended Splash Engine to support native 8-bit Qwen3.8-27B on Apple Silicon, achieving 37-55 tok/s without quantization degradation and scaling up to 256k context.
SwitchSD is an adaptive framework for speculative decoding in LLMs that uses intrinsic model signals to dynamically switch between neural drafting and context-based copying, achieving up to 15% throughput gains over baselines like EAGLE3.
Zarya is a hybrid language model that jointly optimizes autoregressive and masked diffusion objectives for flexible training and dual-mode inference, with publicly released models in sizes 0.6B, 1.7B, and 4B.
ByteShape releases full ShapeLearn quantized versions of the Qwen 3.8 27B model in GGUF format, with benchmarking showing improvements in quality-speed frontier and support for speculative decoding.
A talk about llama.cpp speculative decoding methods (MTP, dflash, dspark) at the dotAI conference will have its replay available soon.
Intel releases OpenVINO 2026.4 with expanded AI model support, performance enhancements like multi-token prediction, and new features for profiling and inference across CPUs, GPUs, and NPUs.
This paper constructs a cost-quality-latency Pareto atlas for LLM inference optimizations, using a calibrated simulator to evaluate configurations and combinations across different hardware and regimes.
FlexEE introduces a self-speculative and KV-cache-compatible early exiting framework for efficient LLM inference in offloading deployments, achieving significant speedups on Llama models with minimal accuracy degradation.
The author optimized DeepSeek V4.1 Flash for Apple M3 Ultra, achieving up to 40 t/s decode speed with DSpark speculative decoding while maintaining byte-identical accuracy to the upstream model.
R9V update delivers approximately 100 tokens per second on Qwen3.8 Flash Next with IQ4_XS on dual AMD R9700 GPUs, adds support for Q4_K_XL with 50 tok/s, fixes crashes, and enhances diagnostics.
This paper investigates the losslessness of Orthrus, a hybrid autoregressive-diffusion model for inference acceleration, finding that it requires high numerical precision (FP32) for exact trajectory matching, while BF16 divergence does not impair downstream performance.
A draft model for Qwen 3.8 27B is highly effective on 16 GB GPUs, achieving around 60 tokens per second on an RX 9070 XT and offering better VRAM efficiency than built-in MTP.
The Qwen 3.8 Flash Next model achieved a 100% score on the CUDA.fast benchmark with a 117.8% composite speed increase on DGX Spark, utilizing speculative decoding.
Osprey introduces a target-agnostic pre-training method for drafters in speculative decoding, improving efficiency by bootstrapping from off-the-shelf models and adapting with minimal target-specific work, achieving significant acceptance rate improvements across multiple LLMs.
This paper introduces X-CoSD, a communication-efficient cross-vocabulary collaborative speculative decoding framework that optimizes distributed LLM inference by splitting residual resampling to reduce overhead while preserving server LLM quality.
The author explains how speculative decoding affects coding agent speed, with higher acceptance rates on boilerplate code leading to faster typing, and discusses other factors like cache misses that impact performance.