Tag
Deepseek has soft retired its V4 Pro model, indicating that the AI model is being phased out or deprecated.
The community discovered that the open-source model architecture of V4-Pro-0813 uploaded by DeepSeek to Hugging Face was actually V4 Flash. The repository was taken down and then re-uploaded, and the hashes and sizes of some safetensors weight files also changed, suspected to be a release/deployment error.
DeepSeek launches V4-Pro and V4-Flash with flexible reasoning effort, native OpenAI Responses API support, and optimized agent workflows for Codex, available via API and app/web.
DeepSeek has quietly released V4 Pro (0813), with API documentation showing support for the Responses API format and integration with Codex, enabling use of deepseek-v4-flash/pro models.
DeepSeek v4 PRO, a 1.6 trillion parameter model, is running via SSD streaming on a 128GB MacBook m5 max, demonstrating local inference of a massive model.
DeepSeek has made its V4-Pro API price cut of 75% permanent, with per-million cached input tokens at just $0.003625 and output tokens at $0.87, about 34 times cheaper than OpenAI's GPT-5.5. The model has 1.6 trillion parameters but requires only 49 billion active parameters, supports a 1-million-token context, and leads in coding and reasoning benchmark tests.
DeepSeek makes the discount for DeepSeek-V4-Pro permanent, extending it until May 31, 2026.
DeepSeek has made the 75% discount on V4 Pro API pricing permanent, reducing input/output token costs significantly.