Up to 3.2x Faster Inference with LFM2.5-DSpark
The release introduces DSpark draft model checkpoints for three LFM2.5 models (1.2B-Instruct, 2.6B, and 8B-A1B), adding a speculative decoding path that trades a minimal memory increase for a large decoding speedup. In standard LLM inference, the decode phase is memory-bound—most latency comes from streaming weights from DRAM to SRAM, not computation. Speculative decoding uses a lightweight draft model to propose candidate tokens, which the target model then verifies in a single forward pass, sharing the weight-loading cost across all verified tokens.
DSpark combines three components: a DFlash-style parallel backbone that produces hidden states for all draft tokens in one forward pass conditioned on the target model's context features; a lightweight sequential Markov-chain head that adds inter-token dependency to raise acceptance rates at later positions; and a confidence-scheduled verifier that predicts each token's survival probability and prunes low-confidence suffixes when verification costs more than it saves. The draft models were trained with a larger and more diverse data mix covering SFT, chat, code, and function-calling data, using attention-only architectures with 5 layers and a block of 9. Each model was trained for 15 epochs, selecting the epoch with the highest acceptance rate rather than the lowest loss. The resulting draft models are small: LFM2.5-1.2B-Instruct has 295.7M total parameters, while the 2.6B and 8B-A1B drafts each have 327.7M.
Measured speedups are up to 3.18x throughput improvement on a GPU and up to 2.87x on-device, and the DSpark path cuts function-calling latency by 57% on average for LFM2.5-2.6B, supporting on-device agentic inference. Quality is preserved exactly: under greedy decoding, a draft token is accepted only if it matches the target model's distribution, and rejected tokens are replaced by the target model's own token, making the emitted sequence identical to baseline greedy by construction. Day-one support for llama.cpp and SGLang is open-sourced upstream.