Model Releases Hugging Face Blog

Deploy local agents everywhere with LFM2.5-2.6B

LFM2.5-2.6Bon-device agentsagentic reinforcement learningedge AI

LFM2.5-2.6B is a 2.6B-parameter model built to power capable agents entirely on-device, supporting tool calling and multi-step workflows while remaining small and fast enough for everyday hardware, from laptops to phones. This enables developers to deploy agents everywhere, keep data on the device, and avoid cloud inference costs.

The model was pre-trained on ~34T tokens, with a mid-training phase extending the context window to 128K. Post-training turns the base model into an agent in four stages: two rounds of SFT weighted heavily toward agentic data; teacher specialization with one specialist per domain; multi-domain on-policy distillation (MOPD) into a single student; and agentic reinforcement learning (Agentic RL) running multi-turn RL inside real agent harnesses. The RL pipeline separates training, rollout generation, and environment execution via a Training Engine, Rollout Engine, RL framework, and a Sandbox Service with a Blackbox Harness and Harness Proxy.

In benchmarks against models up to ~4x its size, LFM2.5-2.6B competes with and often beats larger models. Examples include AIME25: 51.87 (vs. 26.33 for gemma-4-E2B-it and 56.07 for Qwen3.5-9B), Multi-IF: 80.07 (highest of the group), BFCLv4: 56.88, and τ³-Bench Banking: 5.67. It achieves 220 tok/s on an Apple M5 Max and 113 tok/s on an AMD Ryzen CPU, all in under 2.5 GB of memory.

LFM2.5-2.6B lets developers deploy agents on everyday hardware, keep data private, and scale usage without a cloud inference bill, making it a strong option for edge AI applications.

Read original →

← Back to home