DeepSeek v4.1 Flash

Hacker News Top Models

Summary

DeepSeek has introduced DeepSeek-V4.1-Flash, a new AI model designed for enhanced capability, faster inference, native visual understanding, and scalability as part of their latest architecture family.

🚀 Introducing DeepSeek-V4.1-Flash: smarter, faster, more efficient. 🔹 Introducing the smallest model in our new architecture family, with native visual understanding. 🔹 Designed for greater capability, faster inference, higher throughput, and scaling to larger models. 1/6 https://t.co/wxJGiyX56o
Original Article
View Cached Full Text

Cached at: 09/10/26, 08:35 AM

🚀 Introducing DeepSeek-V4.1-Flash: smarter, faster, more efficient.

🔹 Introducing the smallest model in our new architecture family, with native visual understanding. 🔹 Designed for greater capability, faster inference, higher throughput, and scaling to larger models.

1/6

🧠 Asymmetric architecture. More intelligence, less cost.

🔹 552B-parameter MoE. 🔹 New Causal Encoder–Decoder architecture: just 8B active parameters for input, 16B for output. 🔹 New pre-training methods + larger-scale RL post-training deliver benchmark results ahead of flagship models, including DeepSeek-V4-Pro.

2/6

💾 Smaller KV cache. Bigger savings.

Compared with the previous generation, V4.1-Flash’s KV cache needs just: 🔹 1/4 the HBM 🔹 1/8 the SSD storage Cache-hit charges often account for a large share of agent costs. Compressing the cache cuts those costs significantly.

3/6

⚡ V4.1-Flash is now live on the DeepSeek API with native multimodal support.

Set your model to deepseek-flash.

🔹 V4-Flash & V4-Flash-Vision-Exp are retired. For compatibility, deepseek-v4-flash and deepseek-v4-flash-vision-exp temporarily route to V4.1-Flash. 🔹 Tests by multiple parties put V4.1-Flash ahead of V4-Pro on performance, cost, speed & total runtime. We’re phasing out V4-Pro. 🔹 Starting at 04:00 UTC on Sept 14, 2026, all deepseek-v4-pro requests will route to V4.1-Flash at V4.1-Flash rates. This will continue until V4.1-Pro launches.

🤝 Official partners @WorkBuddy_AI (including Codebuddy) & @opencode now fully support V4.1-Flash. Try it today!

4/6

💰 More efficient architecture. Lower API prices.

V4.1-Flash lets us serve more users at a lower cost. We’re passing the savings on to you.

🔹 Peak/off-peak pricing continues to balance demand. 🔹 Off-peak rates are 50% of peak rates. Schedule flexible workloads off-peak to save. 🔹 New pricing takes effect at 04:00 UTC on Sept 10, 2026.

5/6

🌐 Supporting open source. Expanding deployment options.

We’ll work closely with the open-source community on V4.1-Flash inference support and explore more deployment options. Planning a large-scale deployment with 2,000 GPUs + a storage cluster? Let’s talk.

🔹 Model: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash… 🔹 Paper: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/blob/main/DeepSeek_V41_Tech_Report.pdf…

6/6

Similar Articles

DeepSeek V4 Flash Vision is now live !

Reddit r/ArtificialInteligence

DeepSeek has released vision capabilities for its V4 Flash AI model, providing a cheaper inference option through DeepInfra compared to the official API.

deepseek-ai/DeepSeek-V4.1-Flash · Hugging Face

Reddit r/LocalLLaMA

The repository provides prompt encoding and a minimal PyTorch inference implementation for the DeepSeek-V4.1-Flash AI model, including components like vision encoder, MoE, and Hyper-Connections under an MIT License.

DeepSeek-V4-Flash-0731

Product Hunt

DeepSeek announces DeepSeek-V4-Flash-0731, a frontier agent intelligence model positioned as offering advanced capabilities at Flash-level pricing.

deepseek-ai/DeepSeek-V4-Flash

Hugging Face Models Trending

DeepSeek releases DeepSeek-V4-Flash and DeepSeek-V4-Pro, new MoE language models supporting 1 million token contexts with improved efficiency and performance.

DeepSeek-V4-Flash-Vision-Exp

Reddit r/LocalLLaMA

DeepSeek-V4-Flash-Vision-Exp is an experimental or updated AI model from DeepSeek focusing on vision capabilities.