Liquid AI has released LFM2.5-2.6B, a model that runs entirely on local hardware. It plans tasks, calls external tools, and manages multi-step workflows on phones, laptops, and PCs without sending data to the cloud. The system contains 2.69 billion parameters, supports a 131,072-token context window, and uses a vocabulary of 128,000 tokens. Development consumed roughly 34 trillion tokens during pre-training. Two versions are available: LFM2.5-2.6B-Base for fine-tuning and the post-trained LFM2.5-2.6B for agent workloads. Because inference happens locally, every run incurs near-zero marginal cost and keeps data private. The company reports that tool-use and instruction-following scores match models nearly four times larger.
In this article
Is it deployable
The answer is yes. Both checkpoints are public on Hugging Face under the lfm1.0 license. Weights ship in native GGUF, MLX, and ONNX formats, with day-one support in llama.cpp, vLLM, SGLang, and LM Studio.
- Which companies: Solo developers and startups can pilot on hardware they already own. The model decodes at 220 tokens/s on an M5 Max in under 2.5 GB. Mid-market teams can self-host on one GPU: a single NVIDIA H100 SXM5 serves roughly 1.3 billion tokens per day. Enterprises and OEMs can push the same weights to device fleets through GGUF and ONNX. Fine-tuning is available via LoRA with TRL and Unsloth.
- Which industries: Liquid AI targets automotive, consumer electronics, industrial robotics, healthcare, financial services, e-commerce, and defense. Regulated and air-gapped settings benefit most, since no prompt reaches a third-party API.
- Applications: Liquid AI recommends agentic workloads, tool use, data extraction, RAG, and long-context workflows. Practical builds include on-device assistants, offline document triage over 128K inputs, form and invoice extraction, robotics command parsing, and background agents that run continuously without per-token cost. Liquid AI explicitly does not recommend the model for agentic coding or knowledge-heavy tasks.
Architecture and training budget
LFM2.5-2.6B has 2.69 billion total parameters across 30 layers. The stack is 22 double-gated short convolution blocks plus 8 grouped-query attention blocks. Vocabulary size is 128,000 and context length is 131,072 tokens. Pre-training used approximately 34 trillion tokens.
Liquid AI doubled the vocabulary to 128K by extending the existing tokenizer in place rather than retraining from scratch. A dedicated mid-training phase extends context to 128K. The model covers 16 languages and is text-only.
Four-stage post-training
The base checkpoint becomes an agent through four stages.
- First, two consecutive supervised fine-tuning rounds, with an SFT mix roughly seven times the size used for LFM2.5-8B-A1B.
- Second, teacher specialization: one expert per domain, trained with reinforcement learning with verifiable rewards.
- Third, multi-domain on-policy distillation, where the student rolls out under its own policy and each prompt routes to its domain teacher.
- Fourth, agentic reinforcement learning with GRPO inside real harnesses, including Hermes Agent and OpenClaw.
Benchmarks
Liquid AI compared LFM2.5-2.6B against gemma-4-E2B-it (5.1B), gemma-4-E4B-it (8B), Qwen3.5-4B (4.7B) and Qwen3.5-9B (9.7B).
| Benchmark | LFM2.5-2.6B | gemma-4-E4B-it | Qwen3.5-9B |
|---|---|---|---|
| ToolSandbox | 77.83 | 65.00 | 76.44 |
| Multi-IF | 80.07 | 77.35 | 62.55 |
| IFStruct | 85.49 | 76.65 | 78.50 |
| IFBench | 59.17 | 39.24 | 56.47 |
| BFCLv4 | 56.88 | 46.39 | 60.13 |
It leads every instruction-following benchmark reported and nearly every tool use benchmark, trailing Qwen3.5-9B only on BFCLv4. Coding is where larger models keep an edge: LiveCodeBenchv6 is 59.41 versus 69.86 for Qwen3.5-9B.
Key takeaways
- 2.69 billion params, 30 layers (22 short-conv + 8 GQA), 128K context, ~34T training tokens.
- Beats gemma-4-E4B-it and Qwen3.5-9B on ToolSandbox, Multi-IF and IFStruct.
- 220 tok/s on M5 Max, 30 tok/s on phone, under 2.5 GB memory.
- Open weights under lfm1.0, with GGUF, MLX and ONNX from day one.
What it means
Users who need private processing now have a model that fits on a single consumer GPU or even a phone. The 220 tokens/s decode speed on an M5 Max means real-time interaction without cloud latency. Teams handling sensitive data or operating in air-gapped environments can deploy this agent without risking prompt leakage. The open weights and native format support lower the barrier for integration into existing stacks.




