speculative-decoding

Tag

Cards List
#speculative-decoding

@Darkolorin: Today we are releasing our speculative decoding implementation in Uzu. Biggest release since inception of our lab. Init…

X AI KOLs Timeline ↗ · 2026-09-03 Cached

Release of a speculative decoding implementation in Uzu, initially supporting Qwen3.6 27B with upcoming support for Qwen3.8 27B and Muse Glimmer.

0 favorites 0 likes
#speculative-decoding

@ViC305: I DID IT!! DeepSeek-V4-Flash-Vision EXL3 MixedK is now running VISION + DSpark speculative decoding together on ONE DGX…

X AI KOLs Timeline ↗ · 2026-09-03 Cached

User @ViC305 successfully runs DeepSeek-V4-Flash-Vision with EXL3 MixedK and DSpark speculative decoding on a single DGX Spark, achieving improved performance and fixing technical issues for multimodal AI deployment.

0 favorites 0 likes
#speculative-decoding

Running a 2-model literary book-translation pipeline on 2x Tesla P40: gemma-4-26B-A4B at ~40 tok/s + Qwen3.6-35B-A3B at 50-70 tok/s with MTP spec decode — full llama-server flags inside

Reddit r/LocalLLaMA ↗ · 2026-09-02

The article details an open-source tool 'Sunny Narrator' for translating fiction books using a two-model AI pipeline on Tesla P40 GPUs, with optimized llama-server configurations achieving 40-70 tokens per second via MTP speculative decoding.

0 favorites 0 likes
#speculative-decoding

Vision Is Not Overhead: One-Pass Block Drafting for Lossless Speculative Decoding in Vision-Language Models

arXiv cs.AI ↗ · 2026-09-02 Cached

The paper introduces GLANCE, a one-pass block drafting method for lossless speculative decoding in vision-language models, achieving up to 2.93x faster generation without changing output.

0 favorites 0 likes
#speculative-decoding

@akshay_pachaar: https://x.com/akshay_pachaar/status/2094765529231929361

X AI KOLs Following ↗ · 2026-09-01 Cached

This article is a practitioner's guide to running local AI models for agent work, highlighting hardware trade-offs and introducing Magnitude, an open-source inference server that simplifies configuration for optimized performance.

0 favorites 0 likes
#speculative-decoding

A rigor-matched audit of periodic-step layer skipping for efficient llm inference: conflayers versus swift, with a supplemental analysis of trained routing alternatives

arXiv cs.CL ↗ · 2026-09-01 Cached

This paper presents a rigor-matched audit comparing periodic-step layer-skipping methods like ConfLayers and SWIFT for efficient LLM inference, and analyzes trained routing alternatives, finding SWIFT superior in accuracy and true inference speed.

0 favorites 0 likes
#speculative-decoding

Verification-Aware Training for Speculative Decoding

Hugging Face Daily Papers ↗ · 2026-08-31 Cached

Verification-Aware Training (VAT) improves draft models for speculative decoding by simulating sequential verification during training and adapting loss weights to acceptance patterns, leading to enhanced acceptance length and inference speedup.

0 favorites 0 likes
#speculative-decoding

@Oluwaphilemon1: Qwen3.8-27B at 56 tok/s on a 9-year-old GPU. Let that sink in. The GPU? NVIDIA V100 32GB. A card that launched at aroun…

X AI KOLs Timeline ↗ · 2026-08-29 Cached

Achieves 56 tokens per second inference speed for the Qwen3.8-27B model on an NVIDIA V100 GPU, demonstrating cost-effective local AI deployment on older hardware using speculative decoding techniques.

0 favorites 0 likes
#speculative-decoding

@pochenai: Inspired by @percyliang's CS336, I built a static performance model for LLM inference — no dynamic batching/chunking, j…

X AI KOLs Following ↗ · 2026-08-28 Cached

Inspired by CS336, a static performance model for LLM inference provides analytical bounds for VRAM, time-to-first-token, and throughput, covering various configurations and calibrated against public benchmarks.

0 favorites 0 likes
#speculative-decoding

Tencent/Hy4-preview 770B-A49B weight dropped

Reddit r/LocalLLaMA ↗ · 2026-08-28 Cached

Tencent releases Hy4-preview, a new Mixture-of-Experts flagship AI model with 770B total parameters and 49B activated per token, featuring advanced techniques like Gated DeepSeek Sparse Attention and identity Hyper-Connections.

0 favorites 0 likes
#speculative-decoding

TreeGraft: Adaptive Multi-Drafter Grafting for Tree-Based Speculative Decoding

arXiv cs.CL ↗ · 2026-08-28 Cached

TreeGraft introduces a multi-drafter framework for tree-based speculative decoding, optimizing draft tree quality with adaptive scheduling to achieve significant inference speedups over single-drafter methods.

0 favorites 0 likes
#speculative-decoding

Dflash 2 speculative decoding by SamuelOliveirads · Pull Request #2345 · ikawrakow/ik_llama.cpp

Reddit r/LocalLLaMA ↗ · 2026-08-27 Cached

This pull request introduces Dflash 2 speculative decoding and other performance improvements to ik_llama.cpp, a fork of llama.cpp focused on enhancing CPU inference and quantization techniques for LLMs.

0 favorites 0 likes
#speculative-decoding

Getting Qwen3.8-27B with decent speed on my 4080 with 16Gb card

Reddit r/LocalLLaMA ↗ · 2026-08-26

This guide explains how to configure DFlash2 speculative decoding with llama.cpp to achieve up to 1.72x faster inference for the Qwen3.8-27B model on an RTX 4080 GPU with 16GB VRAM.

0 favorites 0 likes
#speculative-decoding

Is Multimodal Speculative Decoding Ready for Diffusion-Based Parallel Drafting? A Survey and Empirical Diagnosis

arXiv cs.AI ↗ · 2026-08-24 Cached

This paper surveys and empirically diagnoses the readiness of multimodal speculative decoding for diffusion-based parallel drafting methods, analyzing various multimodal architectures and providing comparative evaluations.

0 favorites 0 likes
#speculative-decoding

LiLiCorr: Lightweight Likelihood Correlation of Parallel Drafts for Speculative Decoding

arXiv cs.CL ↗ · 2026-08-24 Cached

LiLiCorr is a lightweight likelihood-based model that improves speculative decoding by correlating per-position marginal distributions to enhance token coherence, increasing acceptance length and throughput in language model inference.

0 favorites 0 likes
#speculative-decoding

GRAFT: Adaptive DLM-Based Draft Tree Construction with Target-Distilled Edge Scoring

arXiv cs.CL ↗ · 2026-08-24 Cached

GRAFT introduces a draft-tree construction framework for diffusion language model-based speculative decoding, optimizing edge selection and budget allocation to achieve 2.13×–6.36× speedup over autoregressive decoding with low overhead.

0 favorites 0 likes
#speculative-decoding

Self-Speculation for Faster Reasoning Models

arXiv cs.CL ↗ · 2026-08-24 Cached

This paper introduces SSR, a training-free self-speculative decoding method that leverages chain-of-thought to accelerate reasoning in large language models, achieving up to 24.1% latency reduction on structured generation tasks.

0 favorites 0 likes
#speculative-decoding

28 TPS on Qwen2.5-7B across two separate cloud regions over public WAN using speculative decoding + CUDA Graphs [P]

Reddit r/MachineLearning ↗ · 2026-08-23

ShardFlow, a distributed LLM inference framework, achieves 28 TPS on Qwen2.5-7B across cloud regions by using speculative decoding and CUDA Graphs to mitigate WAN latency.

0 favorites 0 likes
#speculative-decoding

Llama.cpp DSpark PC Tree Fork (up to 3%-29.5% faster!)

Reddit r/LocalLLaMA ↗ · 2026-08-21

An implementation of DSpark PC Tree in a llama.cpp fork shows performance gains up to 29.5% faster inference on Qwen 3.0 models based on benchmark results.

0 favorites 0 likes
#speculative-decoding

@Prince_Canuma: LFM2.5 DSpark by @liquidai is coming to mlx-vlm in v0.6.16 Exact speculative decoding on M5 Max, delivering up to 3.7× …

X AI KOLs Following ↗ · 2026-08-21 Cached

LFM2.5 DSpark by liquidai is integrated into mlx-vlm v0.6.16, enabling up to 3.7× faster speculative decoding on M5 Max with zero output drift for on-device VLM inference.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback