Model Releases arXiv cs.CL

DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression

DeepSeekMoEKV cache compressionlong context

The paper frames the problem around long-horizon agents, which make model workloads increasingly input-heavy. Although prior work has substantially reduced the cost of long-context computation, prefill remains computationally expensive, and large KV caches continue to strain HBM and SSD capacity and data-transfer bandwidth. Together, these compute, storage, and bandwidth demands are described as the primary bottleneck to further lowering deployment costs.

To address this, the authors introduce DeepSeek-V4.1-Flash, a multimodal Mixture-of-Experts (MoE) model with 552B backbone parameters and support for contexts of up to one million tokens. Its Causal Encoder-Decoder (CED) architecture activates 16B parameters per token during decode but only 8B parameters during prefill, substantially improving cost efficiency for agentic workloads. To push the limits of KV cache compression, the model combines cross-layer KV cache reuse in Compressed Sparse Attention 2 (CSA2) with FP4 KV caching.

These designs reduce the model's global KV cache footprint (always in HBM) to 890 bytes per token, roughly 1/4 of the corresponding footprint of DeepSeek-V4-Flash. Additionally, a dedicated deployment optimization called SWA Bounded Replay reduces the persistent KV cache footprint (always on SSD or in host memory) to roughly 1/8 of that of DeepSeek-V4-Flash. Despite the much smaller KV cache footprint, the model delivers substantially better performance than the baseline.

The work also streamlines the DeepSeek-V4 architecture and introduces several efficient architectural extensions. DeepSeek-V4.1-Flash was pretrained on a multimodal corpus comprising 45T tokens and then given comprehensive post-training, yielding strong performance across diverse text-based and multimodal agentic scenarios. Model checkpoints are available at https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash.

Read original →

← Back to home