DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression

Hugging Face Daily Papers Papers

Summary

Introduces DeepSeek-V4.1-Flash, a multimodal Mixture-of-Experts model with 552B parameters, featuring advanced KV cache compression techniques to reduce deployment costs and improve efficiency for long-context agent workloads.

The widespread adoption of long-horizon agents has made model workloads increasingly input-heavy. Although prior work has substantially reduced the cost of long-context computation, prefill remains computationally expensive, and large KV caches continue to strain HBM and SSD capacity and data-transfer bandwidth. Together, these compute, storage, and bandwidth demands constitute the primary bottleneck to further lowering deployment costs. To address this challenge, we introduce DeepSeek-V4.1-Flash, a multimodal Mixture-of-Experts (MoE) model with 552B backbone parameters and support for contexts of up to one million tokens. With its Causal Encoder-Decoder (CED) architecture, the model activates 16B parameters per token during decode but only 8B parameters during prefill, substantially improving cost efficiency for agentic workloads. To push the limits of KV cache compression, DeepSeek-V4.1-Flash combines cross-layer KV cache reuse in Compressed Sparse Attention 2 (CSA2) with FP4 KV caching. These designs reduce its global KV cache footprint (always in HBM) to 890 bytes per token, roughly 1/4 of the corresponding footprint of DeepSeek-V4-Flash. Further, through a dedicated deployment optimization known as SWA Bounded Replay, DeepSeek-V4.1-Flash reduces its persistent KV cache footprint (always on SSD or in host memory) to roughly 1/8 of that of DeepSeek-V4-Flash. Despite its much smaller KV cache footprint, the model delivers substantially better performance than the baseline. In addition, we streamline the DeepSeek-V4 architecture and introduce several efficient architectural extensions. We pretrain DeepSeek-V4.1-Flash on a multimodal corpus comprising 45T tokens and conduct comprehensive post-training, yielding strong performance across diverse text-based and multimodal agentic scenarios. Model checkpoints are available at https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash.
Original Article
View Cached Full Text

Cached at: 09/18/26, 02:59 AM

Paper page - DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression

Source: https://huggingface.co/papers/2609.19969 Authors:

,

,

,

,

,

,

,

,

,

,

,

,

,

,

,

,

,

,

,

,

Abstract

Thewidespreadadoptionoflong-horizonagentshasmademodelworkloadsincreasinglyinput-heavy.Althoughpriorworkhassubstantiallyreducedthecostoflong-contextcomputation,prefillremainscomputationallyexpensive,andlargeKVcachescontinuetostrainHBMandSSDcapacityanddata-transferbandwidth.Together,thesecompute,storage,andbandwidthdemandsconstitutetheprimarybottlenecktofurtherloweringdeploymentcosts.Toaddressthischallenge,weintroduceDeepSeek-V4.1-Flash,amultimodalMixture-of-Experts(MoE)modelwith552Bbackboneparametersandsupportforcontextsofuptoonemilliontokens.WithitsCausalEncoder-Decoder(CED)architecture,themodelactivates16Bparameterspertokenduringdecodebutonly8Bparametersduringprefill,substantiallyimprovingcostefficiencyforagenticworkloads.TopushthelimitsofKVcachecompression,DeepSeek-V4.1-Flashcombinescross-layerKVcachereuseinCompressedSparseAttention2(CSA2)withFP4KVcaching.ThesedesignsreduceitsglobalKVcachefootprint(alwaysinHBM)to890bytespertoken,roughly1/4ofthecorrespondingfootprintofDeepSeek-V4-Flash.Further,throughadedicateddeploymentoptimizationknownasSWABoundedReplay,DeepSeek-V4.1-FlashreducesitspersistentKVcachefootprint(alwaysonSSDorinhostmemory)toroughly1/8ofthatofDeepSeek-V4-Flash.DespiteitsmuchsmallerKVcachefootprint,themodeldeliverssubstantiallybetterperformancethanthebaseline.Inaddition,westreamlinetheDeepSeek-V4architectureandintroduceseveralefficientarchitecturalextensions.WepretrainDeepSeek-V4.1-Flashonamultimodalcorpuscomprising45Ttokensandconductcomprehensivepost-training,yieldingstrongperformanceacrossdiversetext-basedandmultimodalagenticscenarios.Modelcheckpointsareavailableathttps://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash.

View arXiv pageView PDFAdd to collection

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2609.19969 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2609.19969 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2609.19969 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

FlashMemory DeepSeek-V4 Retriever (GitHub Repo)

TLDR AI

Introduces FlashMemory DeepSeek-V4 Retriever, a lightweight model that sparsifies DeepSeek-V4's CSA KV-cache by predicting which chunks will be attended to next, keeping only ~10-15% on-device while matching full-attention performance.

DeepSeek-V4-Flash-0731

Product Hunt

DeepSeek announces DeepSeek-V4-Flash-0731, a frontier agent intelligence model positioned as offering advanced capabilities at Flash-level pricing.

deepseek-ai/DeepSeek-V4-Flash-DSpark

Hugging Face Models Trending

DeepSeek releases V4 series of Mixture-of-Experts language models (Pro 1.6T/49B activated, Flash 284B/13B activated) supporting one-million-token context with hybrid attention and speculative decoding, claiming best open-source model performance.

DeepSeek v4.1 Flash

Hacker News Top

DeepSeek has introduced DeepSeek-V4.1-Flash, a new AI model designed for enhanced capability, faster inference, native visual understanding, and scalability as part of their latest architecture family.