@ramin_m_h: yesterday we made them more compressed! today we make them faster than ever with speculative decoding! up to 4x decode …
Summary
Liquid AI releases DSpark draft models for their LFM series, incorporating speculative decoding to achieve up to 4x decode speedup on device while maintaining output quality.
View Cached Full Text
Cached at: 08/21/26, 07:05 AM
yesterday we made them more compressed! today we make them faster than ever with speculative decoding!
up to 4x decode speed up on device for our 1.2B, 2.6B and 8B moe.
You gotta try these LFMs for function calling applications on device or latency critical load on the cloud! work of art by our very own @tugot17
enjoy
Liquid AI (@liquidai): Today, we release DSpark draft models for LFM2.5-1.2B-Instruct, LFM2.5-2.6B, and LFM2.5-8B-A1B. These add a speculative decoding path that trades a minimal memory increase for a large decoding speedup without changing output quality.
A lightweight draft model proposes a block of
Similar Articles
@Prince_Canuma: LFM2.5 DSpark by @liquidai is coming to mlx-vlm in v0.6.16 Exact speculative decoding on M5 Max, delivering up to 3.7× …
LFM2.5 DSpark by liquidai is integrated into mlx-vlm v0.6.16, enabling up to 3.7× faster speculative decoding on M5 Max with zero output drift for on-device VLM inference.
@dzhulgakov: DSpark from @deepseek_ai ingeniously integrates many speculative decoding ideas to achieve 1.5x to 5x higher throughput…
DSpark from DeepSeek AI integrates speculative decoding ideas to achieve 1.5x to 5x higher throughput in production systems. This thread explains 10 key ideas from the basics.
@danielhanchen: DeepSeek just released DSpark for V4 Flash & Pro, a new speculative decoding method boosting throughput by 51% to 400%!…
DeepSeek released DSpark, a speculative decoding method that boosts throughput by 51% to 400% for V4 Flash & Pro, along with the open-source DeepSpec codebase for training and evaluating draft models.
@jimmysmith1919: Another nice release today. New draft models for speculative decoding of several of our LFM2.5 models. 1.2B: https://hu…
LiquidAI releases draft models for speculative decoding to accelerate their LFM2.5 models, achieving up to 2× faster inference on H100 and Apple silicon without quality degradation.
Up to 3.2x Faster Inference with LFM2.5-DSpark
Liquid AI releases DSpark draft model checkpoints for the LFM2.5 family, enabling up to 3.2x faster inference on GPUs and devices with minimal quality trade-off, and with day-one support for open-source tools like llama.cpp and SGLang.