DeepSeek V4 Flash GA 在 DeepSWE 上排名与 Sonnet 5 和 Grok 4.5 相同
摘要
DeepSeek 宣布推出 V4 Flash GA,声称其在 DeepSWE 基准测试上与 Sonnet 5 和 Grok 4.5 相当,但该说法尚未得到验证。
来源:https://x.com/deepseek_ai/status/2083084415157022911 和 https://deepswe.datacurve.ai/ 仅为合并数据视图。此系 DeepSeek 的声明,尚未经 DeepSWE 验证。
相似文章
新版DeepSeek V4-Flash在ArtificalAnalysis Index上取得50分,比GLM-5.2和GPT-5.6 Luna低1分
DeepSeek的新V4-Flash模型在ArtificalAnalysis Index上得分为50,仅比GLM-5.2和GPT-5.6 Luna低1分。
DeepSeek-V4 Flash 实际表现如何?
对 DeepSeek-V4 Flash 的性能和能力的评估,评估其实际有效性。
DeepSeek v4 Flash 能力显著提升
DeepSeek V4 Flash 在预览更新中展现出显著的基准测试提升,在智能体编码任务上与 GPT-5.6 Terra 互有胜负。
DeepSeek-V4-Flash-0731 在基准测试中现已大幅超越 DeepSeek-V4-Pro-Preview
DeepSeek 的新 V4-Flash-0731 模型在基准测试中的表现现已远远优于 V4-Pro-Preview,标志着该模型系列取得了显著进步。
DeepSeek V4 flash 0731 在 Agent Arena 排名第21位
DeepSeek V4 flash 0731 在 Agent Arena 中排名第21位,低于 Sonnet 4.6 和 Luna,但用户欣赏其开源性,以获得隐私和控制。