LFM2.5-VL-3B for Better and Faster Vision Capabilities for the Edge
LFM2.5-VL-3B is Liquid AI's latest vision-language model designed to run on user hardware. It understands documents and screens, grounds objects via natural language, and can call tools, while answering directly (rather than reasoning step-by-step) to keep responses fast for real-time and on-device applications. Compared to previous releases, it brings four major improvements: stronger screen/UI understanding across devices, improved grounding and object detection, better reasoning across multiple images, and significantly stronger function calling in both text-only and vision-text scenarios.
The model architecture pairs a SigLIP2 400M NaFlex vision encoder with the same pre-trained backbone as the LFM2.5-2.6B text model. It was pre-trained on roughly 34 trillion tokens, with 4x more vision data than before, drawn from curated and synthetic image-caption, OCR, grounding, and instruction-following sets. To support non-Latin scripts, the vocabulary was doubled to 128K by extending the tokenizer in place rather than retraining from scratch. Post-training happens in two stages: first supervised fine-tuning (SFT) with knowledge distillation from a larger teacher and Antidoom training, then multi-reward reinforcement learning (RL).
In benchmark evaluations across multilingual visual comprehension, instruction following, visual math/science reasoning, document understanding, object detection, multi-image understanding, and screen understanding, LFM2.5-VL-3B leads its size class on real-world image tasks. On digital content, it reads documents, charts, and UI elements well. Notable scores include MMStar 63.3, RealWorldQA 73.1, DocVQA 91.1, ChartQA 81.3, and MathVista 68.5, competitive with or better than larger models like Gemma 4 E4B (8B) and Qwen3.5-4B on several vision tasks.
The model's combination of strong multimodal understanding, fast direct answering, and efficient size makes it suitable for on-device applications like UI automation, document parsing, and tool-using agents running locally. Its extended tokenizer also broadens multilingual usability beyond Latin scripts.