draft-models

Tag

Cards List
#draft-models

Verification-Aware Training for Speculative Decoding

Hugging Face Daily Papers · 2026-08-31 Cached

Verification-Aware Training (VAT) improves draft models for speculative decoding by simulating sequential verification during training and adapting loss weights to acceptance patterns, leading to enhanced acceptance length and inference speedup.

0 favorites 0 likes
#draft-models

@jimmysmith1919: Another nice release today. New draft models for speculative decoding of several of our LFM2.5 models. 1.2B: https://hu…

X AI KOLs Timeline · 2026-08-20 Cached

LiquidAI releases draft models for speculative decoding to accelerate their LFM2.5 models, achieving up to 2× faster inference on H100 and Apple silicon without quality degradation.

0 favorites 0 likes
#draft-models

DeepSeek V4 Flash 0731 on Strix Halo: draft model, n_max sweep, and a launch line that actually helps

Reddit r/LocalLLaMA · 2026-08-18

The article presents benchmark results for DeepSeek V4 Flash 0731 on Strix Halo hardware, showing performance with different draft models and n_max settings, concluding that n_max=3 offers the best speed balance.

0 favorites 0 likes
#draft-models

Speculative Decoding Across Languages

arXiv cs.CL · 2026-06-01 Cached

This paper compares three strategies to improve speculative decoding efficiency for non-English languages, finding that task-specific distillation improves acceptance rates but generalizes poorly, while n-gram draft models offer consistent speed-ups despite lower acceptance rates.

0 favorites 0 likes
#draft-models

Draft-OPD: On-Policy Distillation for Speculative Draft Models

Hugging Face Daily Papers · 2026-05-28 Cached

Draft-OPD introduces on-policy distillation with target-assisted rollouts and error replay to overcome the offline-to-inference mismatch in training draft models for speculative decoding, achieving over 5x lossless acceleration and improving upon EAGLE-3 and DFlash by 23% and 13% respectively.

0 favorites 0 likes
#draft-models

ConFu: Contemplate the Future for Better Speculative Sampling

arXiv cs.CL · 2026-04-20 Cached

ConFu introduces a novel speculative decoding framework that enables draft models to anticipate future generation directions through contemplate tokens and soft prompts, achieving 8-20% improvements in token acceptance rates and generation speed over EAGLE-3 across multiple LLM models.

0 favorites 0 likes
← Back to home

Submit Feedback