Accelerating vision-language models with LFM2.5-VL-DSpark
The release introduces an experimental DSpark draft model for the vision-language model LFM2.5-VL-3B. Like the recently released LFM2.5-DSpark drafter models, it adds a speculative decoding path that trades a minimal increase in memory footprint for a larger speedup without changing output quality. The headline gains are decode speedups up to 3.13x on device and 2.66x on an H100, with end-to-end gains up to 2.62x and 2.27x respectively.
For VLMs, the vision drafter uses the same architecture as the text LFM2.5-DSpark drafters: it captures the target model's hidden states at a fixed set of tapped layers and conditions on them to draft a block of k candidate tokens. Image patches and text tokens are projected into a shared representation before those layers, so the drafter operates on hidden-state vectors of identical dimensionality regardless of input modality. The inference algorithm is therefore unchanged from the text models.
Training and architecture follow the DSpark recipe with a mixture of vision-language SFT data, weighted toward the workloads the model is expected to serve. Based on ablations across 3, 4, and 5 layers, the draft model is a simplified attention-only drafter with 4 layers and a block size of 9. The team ran 10 epochs on the final mixture and measured acceptance after each, which improved with additional training tokens before reaching diminishing returns. At inference time, they recommend a block size of 8 or 9 depending on the hardware. The resulting drafter has approximately 280M parameters and increases the deployed model's parameter count by just 8.9%. Its components are: a 4-layer decoder stack at 193.0M parameters, hidden-state projection at 21.0M, Markov head at 65.5M, norms plus confidence head at 6.4k, for a total of 279.5M.
The DSpark draft model ships with day-one support for llama.cpp, MLX-VLM, and SGLang. Both on-device and GPU inference configurations use a DSpark block size of 8 and are evaluated on six diverse vision-based tasks: general VQA, text VQA, image captioning, chart VQA, complex reasoning, and multi-turn conversation, following the MMSpec benchmark. With MLX on an M5 Max, decoding runs 2.30x to 3.13x faster by task, and end-to-end latency improves by 1.56x to 2.62x. With llama.cpp on an M3 Ultra, decoding improves by 1.57x to 2.14x and end-to-end by 1.30x to 1.77x. On H100, the same drafter delivers 20.4x to 2.66x faster decoding, with end-to-end improvements of 1.64x to 2.27x.
The post also begins discussing limitations of speculation for vision workloads: in LLMs, prefill is mostly compute-bound, and its cost grows (sub)quadratically with prompt length. VLMs add to this because the image first passes through a vision encoder, then the language backbone processes hundreds... The excerpt cuts off at that point.