Skip to content
AI IntelligenceSep 11, 2026AI Intelligence
Article

Deepseek releases V4.1-Flash model with 75% lower KV cache memory needs

Deepseek has released V4.1-Flash, a multimodal model with 552 billion parameters that reduces KV cache memory to a quarter of its predecessor. The model has 16 billion active parameters per token and narrowly outperforms Opus 5 and GPT-5.6 Sol on the DeepSWE coding benchmark. It is available under the MIT license.

Frontier EditorialSource: The Decoder
01

Source Brief

Deepseek releases V4.1-Flash model with 75% lower KV cache memory needs: Deepseek has released V4.1-Flash, a multimodal model with 552 billion parameters that reduces KV cache memory to a quarter of its predecessor. The model has 16 billion active parameters per token and narrowly outperforms Opus 5 and GPT-5.6 Sol on the DeepSWE coding benchmark. It is available under the MIT license.