@karanC_12: Ox Alpha just finished the full DeepSWE run. Final score: ~63% For reference: • DeepSeek V4 Pro → 63% • Grok 4.6 → 65% …
Summary
Ox Alpha, a free model with 1M context, achieved a ~63% score on the DeepSWE benchmark, matching frontier mid-tier models and suggesting potential for local AI agent development.
View Cached Full Text
Cached at: 08/23/26, 03:40 PM
Ox Alpha just finished the full DeepSWE run.
Final score: ~63%
For reference: • DeepSeek V4 Pro → 63% • Grok 4.6 → 65% • Gemini 3.7 Flash → 65%
This is a free, nameless model with 1M context that is matching frontier mid-tier models.
If the GLM-5.x Flash rumors are true and this can run on 1-2 DGX Sparks…
The local agent game just changed.
Similar Articles
A stealth model called Ox-Alpha has been released, outperforming Fable on SWE.
A stealth AI model named Ox-Alpha has been released, reportedly outperforming Fable on SWE benchmarks, and is available for free with features like multi-modal support and zero data retention.
I benchmarked Ox Alpha on SWE-bench Verified Mini (50 tasks): 96% resolved. Now I’m skeptical of myself.
The author benchmarked Ox Alpha on SWE-bench Verified Mini, achieving a 96% resolution rate, but raises concerns about the score's validity due to data contamination, small sample size, and non-comparable baselines.
Someone did an audit on the new DeepSWE, the results aren't pretty
DeepSWE is a new benchmark for evaluating AI coding agents on real-world software engineering tasks from active open-source repositories, comprising 113 tasks across TypeScript, Go, Python, JavaScript, and Rust with isolated environments and program-based verifiers.
DeepSWE benchmarks indicate that DeepSeek v4 Pro only passes 8% of tasks
A discussion about DeepSWE benchmarks showing that DeepSeek v4 Pro passes only 8% of tasks, which is surprisingly low compared to its performance on similar tasks.
@VictorKaiWang1: 95.3% on Terminal Bench 2.1 Deepseek V4 Flash + StateM, 88.8% on TB2.1
Researchers achieve 95.3% accuracy on Terminal-Bench 2.1 using DeepSeek V4 Flash and StateM, matching GPT-5.6 Sol Max performance and exploring agent improvement beyond model scaling.