@ramin_m_h: yesterday we made them more compressed! today we make them faster than ever with speculative decoding! up to 4x decode …

X AI KOLs Timeline Models

Summary

Liquid AI releases DSpark draft models for their LFM series, incorporating speculative decoding to achieve up to 4x decode speedup on device while maintaining output quality.

yesterday we made them more compressed! today we make them faster than ever with speculative decoding! up to 4x decode speed up on device for our 1.2B, 2.6B and 8B moe. You gotta try these LFMs for function calling applications on device or latency critical load on the cloud! work of art by our very own @tugot17 enjoy
Original Article
View Cached Full Text

Cached at: 08/21/26, 07:05 AM

yesterday we made them more compressed! today we make them faster than ever with speculative decoding!

up to 4x decode speed up on device for our 1.2B, 2.6B and 8B moe.

You gotta try these LFMs for function calling applications on device or latency critical load on the cloud! work of art by our very own @tugot17

enjoy

Liquid AI (@liquidai): Today, we release DSpark draft models for LFM2.5-1.2B-Instruct, LFM2.5-2.6B, and LFM2.5-8B-A1B. These add a speculative decoding path that trades a minimal memory increase for a large decoding speedup without changing output quality.

A lightweight draft model proposes a block of

Similar Articles

Up to 3.2x Faster Inference with LFM2.5-DSpark

Hugging Face Blog

Liquid AI releases DSpark draft model checkpoints for the LFM2.5 family, enabling up to 3.2x faster inference on GPUs and devices with minimal quality trade-off, and with day-one support for open-source tools like llama.cpp and SGLang.