DeepSeek v4 Flash 能力显著提升
摘要
DeepSeek V4 Flash 在预览更新中展现出显著的基准测试提升,在智能体编码任务上与 GPT-5.6 Terra 互有胜负。
DeepSeek V4 Flash:预览版 → 2026-07-31
基准测试 预览版 0731 差值
Terminal Bench* 56.9 82.7 +25.8
Toolathlon 51.8 70.3 +18.5
NL2Repo — 54.2 新增
Cybergym — 76.7 新增
DeepSWE — 54.4 新增
Agent Last Exam — 25.2 新增
Automation Bench — 25.1 新增
DSBench-FullStack — 68.7 新增
DSBench-Hard — 59.6 新增
* Terminal Bench 已从 v2.0 → v2.1,因此这一提升并非严格的可比比较。
与 GPT-5.6 Terra 对比
基准测试 GPT-5.6 Terra DeepSeek V4 Flash 优势
Terminal Bench 78.4 82.7 Flash(+4.3)
Toolathlon 53.1 70.3 Flash(+17.2)
DeepSWE 69.6 54.4 Terra(+15.2)
Agents' Last Exam 50.4 25.2 Terra(+25.2)
双方互有胜负,但没有明显赢家……非常有趣!
来源:https://api-docs.deepseek.com/updates/
相似文章
DeepSeek-V4-Flash-0731 在基准测试中现已大幅超越 DeepSeek-V4-Pro-Preview
DeepSeek 的新 V4-Flash-0731 模型在基准测试中的表现现已远远优于 V4-Pro-Preview,标志着该模型系列取得了显著进步。
DeepSeek-V4 Flash 实际表现如何?
对 DeepSeek-V4 Flash 的性能和能力的评估,评估其实际有效性。
新版DeepSeek V4-Flash在ArtificalAnalysis Index上取得50分,比GLM-5.2和GPT-5.6 Luna低1分
DeepSeek的新V4-Flash模型在ArtificalAnalysis Index上得分为50,仅比GLM-5.2和GPT-5.6 Luna低1分。
@cline:DeepSeek 一小时前悄悄更新了其更新日志,带来了新的 V4-Flash 升级。其新的 Terminal-Bench 分数为 82.…
DeepSeek 悄悄更新了其更新日志,带来了 V4-Flash 升级,将其 Terminal-Bench 分数提升至 82.7,比 4 月预览版提高了 25.8 分。目前仅通过 API 提供,开源权重即将发布。
DeepSeek V4 Flash 0731 智能、性能与价格分析
对 DeepSeek V4 Flash 0731 的分析,涵盖其智能、性能以及与其他 AI 模型的定价对比。