LiquidAI LFM2.5-VL-3B: a 3.1B local VLM that beats Gemma-4 E4B — screen understanding 2.5 → 82.2
Summary
LiquidAI released LFM2.5-VL-3B, a 3.1B local vision-language model that outperforms Gemma-4 E4B and demonstrates a significant jump in screen understanding, making on-device AI more practical.
Similar Articles
LFM2.5-VL-3B for Better and Faster Vision Capabilities for the Edge
LiquidAI announces LFM2.5-VL-3B, an efficient vision-language model for edge hardware with improved screen understanding, grounding, multi-image input, and function calling, trained with 4x more vision data and post-training via SFT and RL.
LiquidAI/LFM2.5-VL-3B · Hugging Face
LiquidAI releases LFM2.5-VL-3B, a 3B multimodal model for on-device deployment with improved OCR, grounding, and efficient inference, available in multiple formats including GGUF, ONNX, and MLX.
LiquidAI/LFM2.5-230M
Liquid AI released LFM2.5-230M, a compact 230M-parameter hybrid model optimized for on-device deployment with fast edge inference speeds (213 tok/s on Galaxy S25 Ultra) and built for agentic tasks via reinforcement learning.
@ma_sc_: I've been testing this on many other languages than the 14 officially supported and results have been truly surprising.…
A user shares surprising results testing Liquid AI's new LFM2.5-VL-3B vision-language model across many languages, noting strong visual capabilities but weaker instruction following; Liquid AI announces the model can read screens, documents, and ground objects to coordinates.
Liquid AI releases LFM2.5-8B-A1B
Liquid AI released LFM2.5-8B-A1B, an edge model with a 128K context window, 38T tokens of pre-training, and large-scale reinforcement learning, capable of tool calling and complex tasks while fitting on an entry-level laptop.