Tag
Release of a speculative decoding implementation in Uzu, initially supporting Qwen3.6 27B with upcoming support for Qwen3.8 27B and Muse Glimmer.
User @ViC305 successfully runs DeepSeek-V4-Flash-Vision with EXL3 MixedK and DSpark speculative decoding on a single DGX Spark, achieving improved performance and fixing technical issues for multimodal AI deployment.
The article details an open-source tool 'Sunny Narrator' for translating fiction books using a two-model AI pipeline on Tesla P40 GPUs, with optimized llama-server configurations achieving 40-70 tokens per second via MTP speculative decoding.
The paper introduces GLANCE, a one-pass block drafting method for lossless speculative decoding in vision-language models, achieving up to 2.93x faster generation without changing output.
This article is a practitioner's guide to running local AI models for agent work, highlighting hardware trade-offs and introducing Magnitude, an open-source inference server that simplifies configuration for optimized performance.
This paper presents a rigor-matched audit comparing periodic-step layer-skipping methods like ConfLayers and SWIFT for efficient LLM inference, and analyzes trained routing alternatives, finding SWIFT superior in accuracy and true inference speed.
Verification-Aware Training (VAT) improves draft models for speculative decoding by simulating sequential verification during training and adapting loss weights to acceptance patterns, leading to enhanced acceptance length and inference speedup.
Achieves 56 tokens per second inference speed for the Qwen3.8-27B model on an NVIDIA V100 GPU, demonstrating cost-effective local AI deployment on older hardware using speculative decoding techniques.
Inspired by CS336, a static performance model for LLM inference provides analytical bounds for VRAM, time-to-first-token, and throughput, covering various configurations and calibrated against public benchmarks.
Tencent releases Hy4-preview, a new Mixture-of-Experts flagship AI model with 770B total parameters and 49B activated per token, featuring advanced techniques like Gated DeepSeek Sparse Attention and identity Hyper-Connections.
TreeGraft introduces a multi-drafter framework for tree-based speculative decoding, optimizing draft tree quality with adaptive scheduling to achieve significant inference speedups over single-drafter methods.
This pull request introduces Dflash 2 speculative decoding and other performance improvements to ik_llama.cpp, a fork of llama.cpp focused on enhancing CPU inference and quantization techniques for LLMs.
This guide explains how to configure DFlash2 speculative decoding with llama.cpp to achieve up to 1.72x faster inference for the Qwen3.8-27B model on an RTX 4080 GPU with 16GB VRAM.
This paper surveys and empirically diagnoses the readiness of multimodal speculative decoding for diffusion-based parallel drafting methods, analyzing various multimodal architectures and providing comparative evaluations.
LiLiCorr is a lightweight likelihood-based model that improves speculative decoding by correlating per-position marginal distributions to enhance token coherence, increasing acceptance length and throughput in language model inference.
GRAFT introduces a draft-tree construction framework for diffusion language model-based speculative decoding, optimizing edge selection and budget allocation to achieve 2.13×–6.36× speedup over autoregressive decoding with low overhead.
This paper introduces SSR, a training-free self-speculative decoding method that leverages chain-of-thought to accelerate reasoning in large language models, achieving up to 24.1% latency reduction on structured generation tasks.
ShardFlow, a distributed LLM inference framework, achieves 28 TPS on Qwen2.5-7B across cloud regions by using speculative decoding and CUDA Graphs to mitigate WAN latency.
An implementation of DSpark PC Tree in a llama.cpp fork shows performance gains up to 29.5% faster inference on Qwen 3.0 models based on benchmark results.
LFM2.5 DSpark by liquidai is integrated into mlx-vlm v0.6.16, enabling up to 3.7× faster speculative decoding on M5 Max with zero output drift for on-device VLM inference.