LFM2.5 Q4\_0 Checkpoints from Quantization-Aware Distillation
Liquid AI today released QAD Q4_0 GGUFs, updated 4-bit checkpoints for LFM2.5-230M, LFM2.5-350M, LFM2.5-1.2B-Instruct, and LFM2.5-2.6B. These were trained with Quantization-Aware Distillation (QAD), where a high-precision teacher model is distilled into a quantized student, recovering 97% of the BF16 average accuracy lost to quantization while keeping the same low memory footprint and high throughput of native Q4_0 GGUFs.
The checkpoints were benchmarked against post-training quantization (PTQ) GGUFs on a suite covering GPQA Diamond, MMLU-Pro, IFEval, IFBench, Multi-IF, and BFCLv4, with BF16 GGUF as the in-format ceiling, plus GSM8K for the two smaller models and AIME25 for the two larger ones. Across five repeats, the QAD checkpoints retained 97.1%, 96.5%, 97.4%, and 96.6% of their respective BF16 baseline performance.
On real edge hardware (MacBook Pro, NucBox EVO-X2, Samsung Galaxy S26 Ultra, and Raspberry Pi 5), the 230M and 350M QAD Q4_0 checkpoints match Q5_K_M quality within evaluation variance at 4–33% higher decode throughput, while the 1.2B and 2.6B checkpoints match Q4_K_M quality at 3–14% higher throughput. They also match Unsloth's UD-Q4_K_XL where applicable (230M and 1.2B).
All four QAD GGUFs are available on Hugging Face and can be used with llama.cpp or any GGUF Q4_0 runtime, e.g. `llama-cli -hf LiquidAI/LFM2.5-350M --hf-file LFM2.5-350M-QAD-Q4_0.gguf`. The post also includes citation instructions for the release.