@VictorKaiWang1: 95.3% on Terminal Bench 2.1 Deepseek V4 Flash + StateM, 88.8% on TB2.1
Summary
Researchers achieve 95.3% accuracy on Terminal-Bench 2.1 using DeepSeek V4 Flash and StateM, matching GPT-5.6 Sol Max performance and exploring agent improvement beyond model scaling.
View Cached Full Text
Cached at: 08/18/26, 08:28 AM
95.3% on Terminal Bench 2.1 Deepseek V4 Flash + StateM, 88.8% on TB2.1
Qin Ziheng (@henryqin1997): Is model scaling the only source of agent improvement?
We @henryqin1997 @YaxinLu1997 @VITAGroupUT @VictorKaiWang1 are glad to share our work: We reach 95.3% Raw Accuracy, or a $15 Frontier Run (88.8% with DeepSeek-v4-Flash, matching GPT-5.6 Sol Max), on Terminal-Bench 2.1 via
Similar Articles
DeepSeek v4 Flash has a nice bump in Capability
DeepSeek V4 Flash shows significant benchmark gains in preview updates, trading blows with GPT-5.6 Terra on agentic coding tasks.
DeepSeek V4 Pro beats GPT-5.5 Pro on precision
DeepSeek V4 Pro reportedly outperforms GPT-5.5 Pro on precision, suggesting a significant advancement in model accuracy.
New DeepSeek V4-Flash achieves 50 on ArtificalAnalysis Index, 1 point below GLM-5.2 and GPT-5.6 Luna
DeepSeek's new V4-Flash model scores 50 on the ArtificalAnalysis Index, trailing GLM-5.2 and GPT-5.6 Luna by just 1 point.
Deepseek V4 Flash 2-bit quant is the first model I can run locally that achieves 100% in this SQL benchmark
A user reports that DeepSeek V4 Flash, running as a 2-bit quantized GGUF on dual RTX 3080s, is the first local model to score 100% on a real-world SQL benchmark, matching frontier models like Opus 4.7 and GPT-5.5.
@TheAhmadOsman: Some numbers from running DeepSeek V4 Flash 0731 on a DGX Station
Ahmad Osman shares performance numbers from running DeepSeek V4 Flash 0731 on an NVIDIA DGX Station.