@karanC_12: Ox Alpha just finished the full DeepSWE run. Final score: ~63% For reference: • DeepSeek V4 Pro → 63% • Grok 4.6 → 65% …

X AI KOLs Timeline Models

Summary

Ox Alpha, a free model with 1M context, achieved a ~63% score on the DeepSWE benchmark, matching frontier mid-tier models and suggesting potential for local AI agent development.

Ox Alpha just finished the full DeepSWE run. Final score: ~63% For reference: • DeepSeek V4 Pro → 63% • Grok 4.6 → 65% • Gemini 3.7 Flash → 65% This is a free, nameless model with 1M context that is matching frontier mid-tier models. If the GLM-5.x Flash rumors are true and this can run on 1-2 DGX Sparks… The local agent game just changed.
Original Article
View Cached Full Text

Cached at: 08/23/26, 03:40 PM

Ox Alpha just finished the full DeepSWE run.

Final score: ~63%

For reference: • DeepSeek V4 Pro → 63% • Grok 4.6 → 65% • Gemini 3.7 Flash → 65%

This is a free, nameless model with 1M context that is matching frontier mid-tier models.

If the GLM-5.x Flash rumors are true and this can run on 1-2 DGX Sparks…

The local agent game just changed.

Similar Articles

Someone did an audit on the new DeepSWE, the results aren't pretty

Reddit r/singularity

DeepSWE is a new benchmark for evaluating AI coding agents on real-world software engineering tasks from active open-source repositories, comprising 113 tasks across TypeScript, Go, Python, JavaScript, and Rust with isolated environments and program-based verifiers.