@VictorKaiWang1: 95.3% on Terminal Bench 2.1 Deepseek V4 Flash + StateM, 88.8% on TB2.1

X AI KOLs Timeline Papers

Summary

Researchers achieve 95.3% accuracy on Terminal-Bench 2.1 using DeepSeek V4 Flash and StateM, matching GPT-5.6 Sol Max performance and exploring agent improvement beyond model scaling.

95.3% on Terminal Bench 2.1 Deepseek V4 Flash + StateM, 88.8% on TB2.1
Original Article
View Cached Full Text

Cached at: 08/18/26, 08:28 AM

95.3% on Terminal Bench 2.1 Deepseek V4 Flash + StateM, 88.8% on TB2.1

Qin Ziheng (@henryqin1997): Is model scaling the only source of agent improvement?

We @henryqin1997 @YaxinLu1997 @VITAGroupUT @VictorKaiWang1 are glad to share our work: We reach 95.3% Raw Accuracy, or a $15 Frontier Run (88.8% with DeepSeek-v4-Flash, matching GPT-5.6 Sol Max), on Terminal-Bench 2.1 via

Similar Articles