LiquidAI/LFM2.5-230M
Summary
Liquid AI released LFM2.5-230M, a compact 230M-parameter hybrid model optimized for on-device deployment with fast edge inference speeds (213 tok/s on Galaxy S25 Ultra) and built for agentic tasks via reinforcement learning.
View Cached Full Text
Cached at: 06/26/26, 05:21 AM
LiquidAI/LFM2.5-230M · Hugging Face
Source: https://huggingface.co/LiquidAI/LFM2.5-230M

LFM2.5 is a family of hybrid models designed foron-device deployment. It builds on the LFM2 architecture with extended pre-training and reinforcement learning.
- Our most compact model yet: 230M parameters that punch above their weight, bringing real capability to the tightest memory and compute budgets.
- Fast edge inference: Best throughput from low-cost CPUs to production GPUs, running at 213 tok/s decode speed on Galaxy S25 Ultra and 42 tok/s on a Raspberry Pi 5.
- Built for agentic tasks: Distilled from LFM2.5-350M and refined with multi-stage reinforcement learning, making it well-suited for tool use and data extraction.
Find more information about LFM2.5-230M in ourblog post.
https://huggingface.co/LiquidAI/LFM2.5-230M#%F0%9F%97%92%EF%B8%8F-model-details🗒️ Model Details
ModelParametersDescriptionLFM2.5-230M-Base230MPre-trained base model for fine-tuningLFM2.5-230M230MGeneral-purpose instruction-tuned model LFM2.5-230M is a general-purpose text-only model with the following features:
- Number of parameters: 230M
- Number of layers: 14 (8 double-gated LIV convolution blocks + 6 GQA blocks)
- Training budget: 19T tokens
- Context length: 32,768 tokens
- Vocabulary size: 65,536
- Knowledge cutoff: Mid-2024
- Languages: English, Arabic, Chinese, French, German, Italian, Japanese, Korean, Portuguese, Spanish
- Generation parameters:-
temperature: 0\.1-top\_k: 50-repetition\_penalty: 1\.05
ModelDescriptionLFM2.5-230MOriginal model checkpoint in native format. Best for fine-tuning or inference with Transformers, vLLM, and SGLang.LFM2.5-230M-GGUFQuantized format for llama.cpp and compatible tools. Optimized for edge inference and local deployment.LFM2.5-230M-ONNXONNX Runtime format for cross-platform deployment.LFM2.5-230M-MLXMLX format for Apple Silicon. Optimized for fast inference on Mac devices. We recommend using it for data extraction and lightweight on-device agentic pipelines. It is not recommended for reasoning-heavy workloads such as advanced math, code generation, or creative writing.
https://huggingface.co/LiquidAI/LFM2.5-230M#chat-templateChat Template
LFM2.5 uses a ChatML-like format. See theChat Template documentationfor details. Example:
<|startoftext|><|im_start|>system
You are a helpful assistant trained by Liquid AI.<|im_end|>
<|im_start|>user
What is C. elegans?<|im_end|>
<|im_start|>assistant
You can usetokenizer\.apply\_chat\_template\(\)to format your messages automatically.
https://huggingface.co/LiquidAI/LFM2.5-230M#tool-useTool Use
LFM2.5 supports function calling in four steps:
- Function definition: Provide the list of tools as a JSON object in the system prompt, or use
tokenizer\.apply\_chat\_template\(\)withtools=\.\.\.. - Function call: By default, LFM2.5 writes Pythonic function calls (a Python list between
<\|tool\_call\_start\|\>and<\|tool\_call\_end\|\>special tokens), as the assistant answer. You can override this behavior by asking the model to output JSON function calls in the system prompt. - Function execution: Execute the call and return the result with the
toolrole. - Final answer: LFM2.5 interprets the tool output and returns a plain-text answer addressing the original prompt.
See theTool Use documentationfor the full guide. Example:
<|startoftext|><|im_start|>system
List of tools: [{"name": "get_candidate_status", "description": "Retrieves the current status of a candidate in the recruitment process", "parameters": {"type": "object", "properties": {"candidate_id": {"type": "string", "description": "Unique identifier for the candidate"}}, "required": ["candidate_id"]}}]<|im_end|>
<|im_start|>user
What is the current status of candidate ID 12345?<|im_end|>
<|im_start|>assistant
<|tool_call_start|>[get_candidate_status(candidate_id="12345")]<|tool_call_end|>Checking the current status of candidate ID 12345.<|im_end|>
<|im_start|>tool
[{"candidate_id": "12345", "status": "Interview Scheduled", "position": "Clinical Research Associate", "date": "2023-11-20"}]<|im_end|>
<|im_start|>assistant
The candidate with ID 12345 is currently in the "Interview Scheduled" stage for the position of Clinical Research Associate, with an interview date set for 2023-11-20.<|im_end|>
https://huggingface.co/LiquidAI/LFM2.5-230M#%F0%9F%8F%83-inference🏃 Inference
LFM2.5 is supported by many inference frameworks. See theInference documentationfor the full list.
Quick start with Transformers (compatible withtransformers\>=5\.0\.0):
from transformers import AutoModelForCausalLM, AutoTokenizer, TextStreamer
model_id = "LiquidAI/LFM2.5-230M"
model = AutoModelForCausalLM.from_pretrained(
model_id,
device_map="auto",
dtype="bfloat16",
# attn_implementation="flash_attention_2" <- uncomment on compatible GPU
)
tokenizer = AutoTokenizer.from_pretrained(model_id)
streamer = TextStreamer(tokenizer, skip_prompt=True, skip_special_tokens=True)
prompt = "What is C. elegans?"
input_ids = tokenizer.apply_chat_template(
[{"role": "user", "content": prompt}],
add_generation_prompt=True,
return_tensors="pt",
tokenize=True,
)["input_ids"].to(model.device)
output = model.generate(
input_ids,
do_sample=True,
temperature=0.1,
top_k=50,
repetition_penalty=1.05,
max_new_tokens=512,
streamer=streamer,
)
https://huggingface.co/LiquidAI/LFM2.5-230M#%F0%9F%94%A7-fine-tuning🔧 Fine-Tuning
We recommend fine-tuning LFM2.5 for your specific use case to achieve the best results.
https://huggingface.co/LiquidAI/LFM2.5-230M#%F0%9F%93%8A-performance📊 Performance
https://huggingface.co/LiquidAI/LFM2.5-230M#benchmarksBenchmarks
ModelGPQA DiamondMMLU-ProIFEvalIFBenchMulti-IFLFM2.5-230M25.4120.2571.7138.4037.70LFM2.5-350M30.6420.0176.9640.6944.92LFM2-350M27.5819.2964.9618.2032.92Granite 4.0-H-350M22.3213.1461.2717.2228.70Granite 4.0-350M25.9112.8453.4815.9824.21Qwen3.5-0.8B (Instruct)27.4137.4259.9422.8741.68Gemma 3 1B IT23.8914.0463.4920.3344.25 ModelCaseReportBenchBFCLv3BFCLv4τ²-Bench Telecomτ²-Bench RetailLFM2.5-230M22.5143.2621.035.2613.68LFM2.5-350M32.4544.1121.8618.8617.84LFM2-350M11.6722.9512.2910.825.56Granite 4.0-H-350M12.4443.0713.2813.746.14Granite 4.0-350M0.8439.5813.732.926.14Qwen3.5-0.8B (Instruct)13.8335.0818.7012.576.14Gemma 3 1B IT2.2816.617.179.366.43
https://huggingface.co/LiquidAI/LFM2.5-230M#cpu-inferenceCPU Inference
https://huggingface.co/LiquidAI/LFM2.5-230M#gpu-inferenceGPU Inference
https://huggingface.co/LiquidAI/LFM2.5-230M#%F0%9F%93%AC-contact📬 Contact
- Got questions or want to connect?Join our Discord community
- If you are interested in custom solutions with edge deployment, please contactour sales team.
https://huggingface.co/LiquidAI/LFM2.5-230M#citationCitation
@article{liquidAI2026230M,
author = {Liquid AI},
title = {LFM2.5-230M: Built to Run Anywhere},
journal = {Liquid AI Blog},
year = {2026},
note = {www.liquid.ai/blog/lfm2-5-230m},
}
@article{liquidai2025lfm2,
title={LFM2 Technical Report},
author={Liquid AI},
journal={arXiv preprint arXiv:2511.23404},
year={2025}
}
Similar Articles
LiquidAI/LFM2.5-2.6B
Liquid AI released LFM2.5-2.6B, a 2.6B-parameter hybrid model optimized for on-device deployment with 128K context, agentic post-training, and fast inference (220 tok/s on Apple M5 Max) under 2.5GB memory.
@liquidai: Introducing LFM2.5-230M: our smallest model yet, built to run fast anywhere (CPUs, NPUs, and GPUs) to enable agentic ta…
Liquid AI releases LFM2.5-230M, a small 230M parameter model optimized for fast inference on CPUs, NPUs, and GPUs, targeting agentic tasks on devices like phones and robots.
Liquid AI releases LFM2.5-8B-A1B
Liquid AI released LFM2.5-8B-A1B, an edge model with a 128K context window, 38T tokens of pre-training, and large-scale reinforcement learning, capable of tool calling and complex tasks while fitting on an entry-level laptop.
LiquidAI/LFM2.5-VL-3B · Hugging Face
LiquidAI releases LFM2.5-VL-3B, a 3B multimodal model for on-device deployment with improved OCR, grounding, and efficient inference, available in multiple formats including GGUF, ONNX, and MLX.
LFM2.5-2.6B: Deploy Agents Everywhere (8 minute read)
Liquid AI releases LFM2.5-2.6B, a compact agentic model designed to run entirely on-device, enabling free inference, low latency, and privacy. The post details its training pipeline including SFT, teacher specialization, distillation, and agentic RL.


