LiquidAI/LFM2.5-230M

Hugging Face Models Trending Models

Summary

Liquid AI released LFM2.5-230M, a compact 230M-parameter hybrid model optimized for on-device deployment with fast edge inference speeds (213 tok/s on Galaxy S25 Ultra) and built for agentic tasks via reinforcement learning.

Task: text-generation Tags: transformers, safetensors, lfm2, text-generation, liquid, lfm2.5, edge, conversational, en, ar, zh, fr, de, ja, ko, es, pt, it, arxiv:2511.23404, base_model:LiquidAI/LFM2.5-230M-Base, base_model:finetune:LiquidAI/LFM2.5-230M-Base, license:other, endpoints_compatible, region:us
Original Article
View Cached Full Text

Cached at: 06/26/26, 05:21 AM

LiquidAI/LFM2.5-230M · Hugging Face

Source: https://huggingface.co/LiquidAI/LFM2.5-230M Liquid AI

LFM2.5 is a family of hybrid models designed foron-device deployment. It builds on the LFM2 architecture with extended pre-training and reinforcement learning.

  • Our most compact model yet: 230M parameters that punch above their weight, bringing real capability to the tightest memory and compute budgets.
  • Fast edge inference: Best throughput from low-cost CPUs to production GPUs, running at 213 tok/s decode speed on Galaxy S25 Ultra and 42 tok/s on a Raspberry Pi 5.
  • Built for agentic tasks: Distilled from LFM2.5-350M and refined with multi-stage reinforcement learning, making it well-suited for tool use and data extraction.

Find more information about LFM2.5-230M in ourblog post.

lfm2_5_230m_benchmarks

https://huggingface.co/LiquidAI/LFM2.5-230M#%F0%9F%97%92%EF%B8%8F-model-details🗒️ Model Details

ModelParametersDescriptionLFM2.5-230M-Base230MPre-trained base model for fine-tuningLFM2.5-230M230MGeneral-purpose instruction-tuned model LFM2.5-230M is a general-purpose text-only model with the following features:

  • Number of parameters: 230M
  • Number of layers: 14 (8 double-gated LIV convolution blocks + 6 GQA blocks)
  • Training budget: 19T tokens
  • Context length: 32,768 tokens
  • Vocabulary size: 65,536
  • Knowledge cutoff: Mid-2024
  • Languages: English, Arabic, Chinese, French, German, Italian, Japanese, Korean, Portuguese, Spanish
  • Generation parameters:- temperature: 0\.1 - top\_k: 50 - repetition\_penalty: 1\.05

ModelDescriptionLFM2.5-230MOriginal model checkpoint in native format. Best for fine-tuning or inference with Transformers, vLLM, and SGLang.LFM2.5-230M-GGUFQuantized format for llama.cpp and compatible tools. Optimized for edge inference and local deployment.LFM2.5-230M-ONNXONNX Runtime format for cross-platform deployment.LFM2.5-230M-MLXMLX format for Apple Silicon. Optimized for fast inference on Mac devices. We recommend using it for data extraction and lightweight on-device agentic pipelines. It is not recommended for reasoning-heavy workloads such as advanced math, code generation, or creative writing.

https://huggingface.co/LiquidAI/LFM2.5-230M#chat-templateChat Template

LFM2.5 uses a ChatML-like format. See theChat Template documentationfor details. Example:

<|startoftext|><|im_start|>system
You are a helpful assistant trained by Liquid AI.<|im_end|>
<|im_start|>user
What is C. elegans?<|im_end|>
<|im_start|>assistant

You can usetokenizer\.apply\_chat\_template\(\)to format your messages automatically.

https://huggingface.co/LiquidAI/LFM2.5-230M#tool-useTool Use

LFM2.5 supports function calling in four steps:

  1. Function definition: Provide the list of tools as a JSON object in the system prompt, or usetokenizer\.apply\_chat\_template\(\)withtools=\.\.\..
  2. Function call: By default, LFM2.5 writes Pythonic function calls (a Python list between<\|tool\_call\_start\|\>and<\|tool\_call\_end\|\>special tokens), as the assistant answer. You can override this behavior by asking the model to output JSON function calls in the system prompt.
  3. Function execution: Execute the call and return the result with thetoolrole.
  4. Final answer: LFM2.5 interprets the tool output and returns a plain-text answer addressing the original prompt.

See theTool Use documentationfor the full guide. Example:

<|startoftext|><|im_start|>system
List of tools: [{"name": "get_candidate_status", "description": "Retrieves the current status of a candidate in the recruitment process", "parameters": {"type": "object", "properties": {"candidate_id": {"type": "string", "description": "Unique identifier for the candidate"}}, "required": ["candidate_id"]}}]<|im_end|>
<|im_start|>user
What is the current status of candidate ID 12345?<|im_end|>
<|im_start|>assistant
<|tool_call_start|>[get_candidate_status(candidate_id="12345")]<|tool_call_end|>Checking the current status of candidate ID 12345.<|im_end|>
<|im_start|>tool
[{"candidate_id": "12345", "status": "Interview Scheduled", "position": "Clinical Research Associate", "date": "2023-11-20"}]<|im_end|>
<|im_start|>assistant
The candidate with ID 12345 is currently in the "Interview Scheduled" stage for the position of Clinical Research Associate, with an interview date set for 2023-11-20.<|im_end|>

https://huggingface.co/LiquidAI/LFM2.5-230M#%F0%9F%8F%83-inference🏃 Inference

LFM2.5 is supported by many inference frameworks. See theInference documentationfor the full list.

Quick start with Transformers (compatible withtransformers\>=5\.0\.0):

from transformers import AutoModelForCausalLM, AutoTokenizer, TextStreamer

model_id = "LiquidAI/LFM2.5-230M"
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    device_map="auto",
    dtype="bfloat16",
#   attn_implementation="flash_attention_2" <- uncomment on compatible GPU
)
tokenizer = AutoTokenizer.from_pretrained(model_id)
streamer = TextStreamer(tokenizer, skip_prompt=True, skip_special_tokens=True)

prompt = "What is C. elegans?"

input_ids = tokenizer.apply_chat_template(
    [{"role": "user", "content": prompt}],
    add_generation_prompt=True,
    return_tensors="pt",
    tokenize=True,
)["input_ids"].to(model.device)

output = model.generate(
    input_ids,
    do_sample=True,
    temperature=0.1,
    top_k=50,
    repetition_penalty=1.05,
    max_new_tokens=512,
    streamer=streamer,
)

https://huggingface.co/LiquidAI/LFM2.5-230M#%F0%9F%94%A7-fine-tuning🔧 Fine-Tuning

We recommend fine-tuning LFM2.5 for your specific use case to achieve the best results.

https://huggingface.co/LiquidAI/LFM2.5-230M#%F0%9F%93%8A-performance📊 Performance

https://huggingface.co/LiquidAI/LFM2.5-230M#benchmarksBenchmarks

ModelGPQA DiamondMMLU-ProIFEvalIFBenchMulti-IFLFM2.5-230M25.4120.2571.7138.4037.70LFM2.5-350M30.6420.0176.9640.6944.92LFM2-350M27.5819.2964.9618.2032.92Granite 4.0-H-350M22.3213.1461.2717.2228.70Granite 4.0-350M25.9112.8453.4815.9824.21Qwen3.5-0.8B (Instruct)27.4137.4259.9422.8741.68Gemma 3 1B IT23.8914.0463.4920.3344.25 ModelCaseReportBenchBFCLv3BFCLv4τ²-Bench Telecomτ²-Bench RetailLFM2.5-230M22.5143.2621.035.2613.68LFM2.5-350M32.4544.1121.8618.8617.84LFM2-350M11.6722.9512.2910.825.56Granite 4.0-H-350M12.4443.0713.2813.746.14Granite 4.0-350M0.8439.5813.732.926.14Qwen3.5-0.8B (Instruct)13.8335.0818.7012.576.14Gemma 3 1B IT2.2816.617.179.366.43

https://huggingface.co/LiquidAI/LFM2.5-230M#cpu-inferenceCPU Inference

image

https://huggingface.co/LiquidAI/LFM2.5-230M#gpu-inferenceGPU Inference

image

https://huggingface.co/LiquidAI/LFM2.5-230M#%F0%9F%93%AC-contact📬 Contact

https://huggingface.co/LiquidAI/LFM2.5-230M#citationCitation

@article{liquidAI2026230M,
  author = {Liquid AI},
  title = {LFM2.5-230M: Built to Run Anywhere},
  journal = {Liquid AI Blog},
  year = {2026},
  note = {www.liquid.ai/blog/lfm2-5-230m},
}
@article{liquidai2025lfm2,
  title={LFM2 Technical Report},
  author={Liquid AI},
  journal={arXiv preprint arXiv:2511.23404},
  year={2025}
}

Similar Articles

LiquidAI/LFM2.5-2.6B

Hugging Face Models Trending

Liquid AI released LFM2.5-2.6B, a 2.6B-parameter hybrid model optimized for on-device deployment with 128K context, agentic post-training, and fast inference (220 tok/s on Apple M5 Max) under 2.5GB memory.

Liquid AI releases LFM2.5-8B-A1B

Reddit r/LocalLLaMA

Liquid AI released LFM2.5-8B-A1B, an edge model with a 128K context window, 38T tokens of pre-training, and large-scale reinforcement learning, capable of tool calling and complex tasks while fitting on an entry-level laptop.

LiquidAI/LFM2.5-VL-3B · Hugging Face

Reddit r/LocalLLaMA

LiquidAI releases LFM2.5-VL-3B, a 3B multimodal model for on-device deployment with improved OCR, grounding, and efficient inference, available in multiple formats including GGUF, ONNX, and MLX.

LFM2.5-2.6B: Deploy Agents Everywhere (8 minute read)

TLDR AI

Liquid AI releases LFM2.5-2.6B, a compact agentic model designed to run entirely on-device, enabling free inference, low latency, and privacy. The post details its training pipeline including SFT, teacher specialization, distillation, and agentic RL.