@VictorKaiWang1:在 Terminal Bench 2.1 上达到 95.3%,DeepSeek V4 Flash + StateM,在 TB2.1 上为 88.8%
摘要
研究人员使用 DeepSeek V4 Flash 和 StateM 在 Terminal-Bench 2.1 上达到了 95.3% 的准确率,匹配了 GPT-5.6 Sol Max 的性能,并探索了超越模型扩展的智能体改进。
查看缓存全文
缓存时间: 2026/08/18 08:28
在Terminal Bench 2.1上达到95.3%的准确率
Deepseek V4 Flash + StateM,于TB2.1上获得88.8%的成绩
秦子恒 (@henryqin1997):
模型规模化是否是智能体改进的唯一路径?我们 @henryqin1997 @YaxinLu1997 @VITAGroupUT @VictorKaiWang1 很高兴分享我们的工作:我们通过(采用DeepSeek-v4-Flash达到88.8%,与GPT-5.6 Sol Max持平)$15前沿方案(Frontier Run)在Terminal-Bench 2.1上实现了95.3%的原始准确率。
相似文章
DeepSeek v4 Flash 能力显著提升
DeepSeek V4 Flash 在预览更新中展现出显著的基准测试提升,在智能体编码任务上与 GPT-5.6 Terra 互有胜负。
DeepSeek V4 Pro 在精确度上击败 GPT-5.5 Pro
据报道,DeepSeek V4 Pro 在精确度上优于 GPT-5.5 Pro,这标志着模型准确性方面的重大进步。
新版DeepSeek V4-Flash在ArtificalAnalysis Index上取得50分,比GLM-5.2和GPT-5.6 Luna低1分
DeepSeek的新V4-Flash模型在ArtificalAnalysis Index上得分为50,仅比GLM-5.2和GPT-5.6 Luna低1分。
Deepseek V4 Flash 2-bit 量化是我能在本地运行且在此 SQL 基准测试中达到 100% 的首个模型
一位用户报告称,DeepSeek V4 Flash 以 2-bit 量化 GGUF 形式在双 RTX 3080 上运行,是首个在真实世界 SQL 基准测试中取得 100% 得分的本地模型,与 Opus 4.7 和 GPT-5.5 等前沿模型持平。
@TheAhmadOsman:在 DGX Station 上运行 DeepSeek V4 Flash 0731 的一些数据
Ahmad Osman 分享了在 NVIDIA DGX Station 上运行 DeepSeek V4 Flash 0731 的性能数据。