LFM2.5-VL-3B for Better and Faster Vision Capabilities for the Edge

Hugging Face Blog 模型

摘要

LiquidAI announces LFM2.5-VL-3B, an efficient vision-language model for edge hardware with improved screen understanding, grounding, multi-image input, and function calling, trained with 4x more vision data and post-training via SFT and RL.

暂无内容
查看原文
查看缓存全文

缓存时间: 2026/08/12 14:20

LFM2.5-VL-3B for Better and Faster Vision Capabilities for the Edge

Source: https://huggingface.co/blog/LiquidAI/lfm2-5-vl-3b Back to Articles

LFM2.5-VL-3Bis our most capable vision-language model you can run on your own hardware. It understands documents and screens alike, grounds objects, and can call tools. It answers directly instead of reasoning, so responses stay fast in real-time and on-device apps.

LFM2.5-VL-3B extends the vision-language capabilities of our previous releases with four major improvements:

  • **Screen/UI understanding:**Strong understanding of digital screens across different devices.
  • **Grounding:**Improved grounding and object detection with natural language queries.
  • **Multi-image input:**Improved reasoning across multiple images.
  • **Function calling:**Significantly stronger at function calling, in text-only and vision-text situations.

lfm2_5_vl_3b_task_group_averages

https://huggingface.co/blog/LiquidAI/lfm2-5-vl-3b#how-we-trained-our-most-capable-vision-language-modelHow we trained our most capable vision-language model

LFM2.5-VL-3B pairs aSigLIP2 400M NaFlex vision encoderwith the same pre-trained backbone as ourLFM2.5-2.6Btext model. It is pre-trained on about 34T tokens, with 4x more vision data than before, drawn from curated and synthetic image-caption, OCR, grounding, and instruction-following sets. To support non-Latin scripts, we doubled the vocabulary to 128K byextending the tokenizer in placerather than retraining from scratch.

Post-training runs in two stages: First is supervised fine-tuning (SFT), with knowledge distillation from a larger teacher andAntidoom training. Second is multi-reward reinforcement learning (RL).

https://huggingface.co/blog/LiquidAI/lfm2-5-vl-3b#benchmark-resultsBenchmark results

We evaluated LFM2.5-VL-3B across both vision and text benchmarks.

Thevision benchmarkscover multilingual visual comprehension, instruction following, visual math and scientific reasoning, document understanding, object detection, multi-image understanding, and screen understanding. LFM2.5-VL-3B leads its size class on real-world image tasks, while also reading digital content well, from documents and charts to on-screen UI elements.

TaskBenchmarkLFM2.5-VL-3B (3.1B)LFM2-VL-3B (3.1B)gemma-4-E2B-it (5.1B)gemma-4-E4B-it (8B)InternVL 3.5 2B (2.4B)InternVL 3.5 4B (4.7B)Qwen3.5-2B (2.3B)Qwen3.5-4B (4.7B)General****MMStar63.357.745.352.957.765.555.159.3MME73.173.054.967.673.681.076.279.5RealWorldQA73.171.160.064.361.667.765.167.1SimpleVQA35.433.027.330.430.533.735.240.7SEED-Bench (image)77.776.671.475.375.476.475.876.1MMBench (dev EN v1.1)81.080.064.271.676.281.173.178.4CountBenchQA87.392.270.480.570.482.583.886.7Multilingual****MMMB83.081.973.380.476.381.575.982.0Multilingual MMBench79.576.362.871.270.976.669.977.0Multimodal IF****MM-IFEval60.651.465.668.247.154.555.463.1STEM****LogicVista37.432.229.534.530.936.234.037.6MathVista (mini)68.562.137.845.256.867.148.763.6MMMU-Pro30.528.726.932.621.322.724.936.0MMMU (val)48.445.641.149.352.060.744.150.3Document, OCR & Chart****ChartQA (test)81.380.443.242.181.786.278.484.2DocVQA (val)91.189.885.787.488.491.892.694.8InfographicVQA (val)70.267.854.460.969.376.973.580.3OCRBench v184.281.770.273.583.982.084.485.6OCRBench v2 (En)47.543.944.448.845.549.147.758.7TextVQA (val)84.383.062.569.076.677.577.381.2Grounding****RefCOCO-avg87.957.167.372.182.988.878.586.6Multi-Image****BLINK61.550.245.252.252.057.248.658.7MuirBench58.334.932.951.845.053.548.262.0Hallucination****HallusionBench47.246.441.849.847.652.149.351.7POPE88.789.284.086.988.088.988.686.0GUI****ScreenSpot-v2 Desktop78.76.028.145.879.982.063.876.3ScreenSpot-v2 Mobile81.27.642.960.386.287.869.781.4ScreenSpot-v2 Web82.22.522.447.679.982.665.977.8**Average****-**69.457.252.059.764.669.463.770.1*All values in the table are normalized to 0–100. Evaluation is done using vLLM 0.26.0 and each model’s recommended generation parameters when available. Non-reasoning mode is used everywhere, and models are prompted to directly answer without reasoning.

We also evaluated LFM2.5-VL-3B ontext-only benchmarksfor instruction following and tool use. Instruction following climbs across the board, and tool use improves sharply. On tool use, LFM2.5-VL-3B is on par with Gemma-4-E2B and Qwen3.5-2B.

TaskBenchmarkLFM2.5-VL-3B (3.1B)LFM2-VL-3B (3.1B)gemma-4-E2B-it (5.1B)gemma-4-E4B-it (8B)InternVL 3.5 2B (2.4B)InternVL 3.5 4B (4.7B)Qwen3.5-2B (2.3B)Qwen3.5-4B (4.7B)Instruction following****IFEval82.372.983.087.932.435.473.686.2IFBench25.820.834.139.224.424.528.933.5Multi-IF59.446.569.477.416.316.953.566.7Tool use & function calling****ToolSandbox59.526.456.561.6N/AN/A47.765.0BFCL V432.520.533.240.0N/AN/A33.953.6*InternVL 3.5 models do not support function-calling.

These results demonstrate that LFM2.5-VL-3B is a strong, general-purpose vision-language model. It covers everyday tasks (captioning, visual question answering, document understanding) and is especially good at grounding objects, reading screens and documents, and calling tools.

https://huggingface.co/blog/LiquidAI/lfm2-5-vl-3b#inference-speed-on-cpu-and-gpuInference speed on CPU and GPU

LFM2.5-VL-3B ships with day-one support across the inference ecosystem, including llama.cpp, MLX, vLLM, SGLang, and ONNX.

**On-device inference.**LFM2.5-VL-3B decodes 228 tokens/s on an M5 Max and 116 tokens/s on a Ryzen AI Max+ 395, and fits in about 3 GB of memory. It even reaches 20 tokens/s on a Galaxy S26 Ultra, so you can run it fully on-device.

lfm2_5_vl_3b_on-device_inference_TTFT

**GPU inference.**LFM2.5-VL-3B keeps latency consistently low and is the fastest on multi-frame inputs.

lfm2_5_vl_3b_ttft

LFM2.5-VL-3B is also the fastest on output throughput out of all models we tested, reaching about 11K tokens per second at high concurrency. That is roughly 2× the larger 4B-class models and ahead of even the smaller 2B-class models, which adds up to nearly 1B output tokens per day on a single H100.

lfm2_5_vl_3b_throughput

https://huggingface.co/blog/LiquidAI/lfm2-5-vl-3b#how-to-use-lfm25-vl-3bHow to use LFM2.5-VL-3B

Reach for LFM2.5-VL-3B when you need on-device intelligence for high-volume workloads.

Install the latest version oftransformers(compatible withtransformers\>=5\.0\.0):

%pip install -q torch torchvision accelerate "transformers>=5.10.1"

Then load and run the model:

import torch
from transformers.image_utils import load_image
from transformers import AutoModelForImageTextToText, AutoProcessor
from IPython.display import display

MODEL_ID = "LiquidAI/LFM2.5-VL-3B" 

processor = AutoProcessor.from_pretrained(MODEL_ID)
model = AutoModelForImageTextToText.from_pretrained(
    MODEL_ID,
    device_map="auto",
    dtype="bfloat16",
)

img_url = "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/coco_sample.png"
input_image = load_image(img_url)
display(input_image)

messages = [
    {
        "role": "user",
        "content": [
            {"type": "image", "image": input_image},
            {"type": "text", "text": "Describe this image in two concise sentences."},
        ],
    }
]

inputs = processor.apply_chat_template(
    messages,
    add_generation_prompt=True,
    tokenize=True,
    return_dict=True,
    return_tensors="pt",
).to(model.device)

with torch.inference_mode():
    outputs = model.generate(
        **inputs,
        do_sample=True,
        temperature=0.2,
        top_k=50,
        repetition_penalty=1.0,
        max_new_tokens=256,
    )

output = processor.batch_decode(outputs[:, inputs["input_ids"].shape[1]:], skip_special_tokens=True)[0]
print(output)

cats_image

Two cats are sleeping on a pink couch with two remote controls.

You can find more hands-on examples on how to use LFM2.5-VL3B for multi-image inputs, grounding, OCR, tool calling, and more inour documentation. Check out ourrelease blogfor video examples.

https://huggingface.co/blog/LiquidAI/lfm2-5-vl-3b#lfm25-vl-3b-demoLFM2.5-VL-3B demo

Check out thisbrowser demo of LFM2.5-VL-3B powering a vision-capable chat interface. It allows you to take or upload multiple images and let the model interact with them, including grounding, OCR, and tool use.

https://huggingface.co/blog/LiquidAI/lfm2-5-vl-3b#get-startedGet Started

LFM2.5-VL-3B is available on Hugging Face today.

With LFM2.5, we’re delivering on our vision of AI that runs anywhere. These models are:

We can’t wait to see what you build.

https://huggingface.co/blog/LiquidAI/lfm2-5-vl-3b#citationCitation

Please cite this article as:

Liquid AI, "LFM2.5-VL-3B: A Better and Faster Vision-Language Model for the Edge", Liquid AI Blog, Aug 2026.

Or use the BibTeX citation:

@article{liquidAI2026VL3B,
  author  = {Liquid AI},
  title   = {LFM2.5-VL-3B: A Better and Faster Vision-Language Model for the Edge},
  journal = {Liquid AI Blog},
  year    = {2026},
  note    = {www.liquid.ai/blog/lfm2-5-vl-3b},
}

相似文章

LiquidAI/LFM2.5-VL-3B · Hugging Face

Reddit r/LocalLLaMA

LiquidAI releases LFM2.5-VL-3B, a 3B multimodal model for on-device deployment with improved OCR, grounding, and efficient inference, available in multiple formats including GGUF, ONNX, and MLX.

Liquid AI 发布 LFM2.5-8B-A1B

Reddit r/LocalLLaMA

Liquid AI 发布了 LFM2.5-8B-A1B,这是一款边缘模型,拥有 128K 上下文窗口、38T 预训练 token 和大规模强化学习,支持工具调用和复杂任务,同时可运行于入门级笔记本电脑。

LiquidAI/LFM2.5-230M

Hugging Face Models Trending

Liquid AI发布了LFM2.5-230M,一款紧凑的230M参数混合模型,针对设备端部署进行了优化,边缘推理速度快(在Galaxy S25 Ultra上达到213 tok/s),并通过强化学习构建,适用于智能体任务。