LFM2.5-VL-3B recognizes Steve from Minecraft running locally on an iPhone 17

Reddit r/LocalLLaMA Models

Summary

Liquid AI released LFM2.5-VL-3B, a 3.1B vision model that runs locally on an iPhone 17 and can recognize objects like a Steve toy from Minecraft, with significantly improved spatial grounding (ScreenSpot-v2 desktop from 6 to 78.7).

Liquid AI put out LFM2.5-VL-3B today, which is a 3.1B vision model that weighs roughly 2GB and fits well on a phone Benchmarks are benchmarks so I tried something sillier. Took a photo of a little Steve toy I have, gave it to the model and asked it what it was looking at It ended up thinking for around 2 minutes and 31 seconds on an iPhone 17, which is a bit too lengthy, but it did end up recognizing Steve and gave a detailed description of him The main diff from the last gen is that it got much better at spotting where things are. ScreenSpot-v2 desktop went from 6 to 78.7. That's why it describes Steve part by part rather than just naming him LFM2.5-VL-3B HF card: https://huggingface.co/LiquidAI/LFM2.5-VL-3B The model was run through atomic.chat mobile app (I'm the founder, so any feedback is welcome)
Original Article

Similar Articles

LiquidAI/LFM2.5-VL-3B · Hugging Face

Reddit r/LocalLLaMA

LiquidAI releases LFM2.5-VL-3B, a 3B multimodal model for on-device deployment with improved OCR, grounding, and efficient inference, available in multiple formats including GGUF, ONNX, and MLX.

LiquidAI/LFM2.5-2.6B

Hugging Face Models Trending

Liquid AI released LFM2.5-2.6B, a 2.6B-parameter hybrid model optimized for on-device deployment with 128K context, agentic post-training, and fast inference (220 tok/s on Apple M5 Max) under 2.5GB memory.

Liquid AI releases LFM2.5-8B-A1B

Reddit r/LocalLLaMA

Liquid AI released LFM2.5-8B-A1B, an edge model with a 128K context window, 38T tokens of pre-training, and large-scale reinforcement learning, capable of tool calling and complex tasks while fitting on an entry-level laptop.