Tag
The benchmark compares DFlash2 and MTP techniques in llama.cpp, showing that DFlash2 offers around 20% faster token generation but reduces available context by 38%.
A developer achieved up to 381 tok/s inference speed on a single RTX 3090 with the Qwen3.8-27B model using optimized techniques like DFlash2 and prefix caching, particularly effective for document-based tasks like RAG and coding assistants.
The Ornith 1.5 35B A3B model's MTP tensors appear to be uninitialized, causing poor speculative decoding performance, and grafting the trained head from Qwen3.6-35B-A3B improves speed by 29%.
LiquidAI releases draft models for speculative decoding to accelerate their LFM2.5 models, achieving up to 2× faster inference on H100 and Apple silicon without quality degradation.
Liquid AI releases DSpark draft models for their LFM series, incorporating speculative decoding to achieve up to 4x decode speedup on device while maintaining output quality.
Liquid AI releases DSpark draft model checkpoints for the LFM2.5 family, enabling up to 3.2x faster inference on GPUs and devices with minimal quality trade-off, and with day-one support for open-source tools like llama.cpp and SGLang.
The author benchmarked llama.cpp flags on a hybrid GPU setup with an RTX 4090 laptop and AMD XTX 7900 eGPU, achieving 70% faster generation, 40% faster prefill, and discovering a bug related to MTP in multi-GPU configurations.
This paper introduces HB-SJD, a batched speculative Jacobi decoding method for visual on-policy distillation that accelerates rollout generation by processing multiple tokens in parallel, reducing training time while preserving generation quality.
The author achieved performance parity between four 2017 Tesla V100 GPUs and a modern RTX 5090 when running the Qwen 3.8 model with NVFP4 precision, using custom software optimizations like the QPN kernel for efficient inference.
DFlash 2 improves speculative decoding by predicting tokens in parallel, achieving over 20% more output per verification pass with minimal latency, and is integrated into major inference engines like SGLang and vLLM.
The article presents benchmark results for DeepSeek V4 Flash 0731 on Strix Halo hardware, showing performance with different draft models and n_max settings, concluding that n_max=3 offers the best speed balance.
DFlash 2 is a block-diffusion draft model for speculative decoding that improves inference speed for the Qwen3.8-27B language model, offering higher acceptance length and throughput in benchmarks.
The article details an experiment achieving 50 tokens per second inference with Qwen3.8-27B at 256K context on a 24GB GPU using Multi-Token Prediction and custom optimizations.
The author tested the Qwen3.8-27B Q8_0 AI model on a ROG Flow Z13 with Ryzen AI Max+ 395, achieving impressive local inference performance in generating a flight simulator using Lemonade Server and llama.cpp with speculative decoding.
This article presents a quantized version of the Qwen3.8-27B model using INT4 AutoRound quantization with working MTP for speculative decoding, achieving significant inference speedups on GPUs.
Introduces DFlash 2, a block-diffusion drafter for speculative decoding with the Qwen3.8-27B model, demonstrating improved acceptance length and throughput in benchmarks.
mlx-dspark v0.10.0 adds support for Qwen3.8-27B on Apple Silicon, providing up to 3x faster inference through speculative decoding with lossless verification.
NInfer adds day-0 support for the Qwen3.8-27B model, achieving around 200 tokens per second on a single RTX 5090 with speculative decoding, and includes engine improvements like concurrent requests and kernel optimizations.
FlashDrive is an algorithm-system co-design framework that cuts the inference latency of vision-language-action models for autonomous driving by 4.7× (from 717 ms to 151 ms on a single GPU) using streaming KV-cache reuse, non-autoregressive diffusion drafting, and adaptive step caching, with negligible accuracy loss.
This paper introduces Decoupled Contrastive Decoding (DCD), which uses an expert-aligned lightweight proposer for speculative decoding while keeping the contrastive signal only in verification, achieving speedups over vanilla contrastive decoding without degrading output distribution.