@JinjingLiang: Codex / Claude / Grok all down. Never imagined `agy` to be so load-bearing...
Summary
The post highlights Gemini 3.8 Flash's superior performance over Opus 5 on DeepSWE-bench, emphasizing its capabilities in coding and agentic tasks, and notes its accessibility through the `agy` harness in Orca.
View Cached Full Text
Cached at: 09/03/26, 06:14 PM
Codex / Claude / Grok all down.
Never imagined agy to be so load-bearing…
Orca ADE (@orca_build): Gemini 3.8 Flash now ranks ahead of Opus 5 on DeepSWE-bench.
It shows particular strength in front-end coding and agentic tasks, while having a fast response time.
You can run it via the
agyharness inside Orca, where it works well paired with complementary models like Fable
Similar Articles
Gemini 3.5 Flash Looks Good For How Fast It Is (8 minute read)
Google released Gemini 3.5 Flash, a hybrid speed model that rivals Opus 4.7 and GPT-5.5 in speed and cost while performing well on agentic and coding benchmarks.
A week running Claude Code, Codex, and Gemini CLI as coding agents on the same repo. Where each one actually breaks.
A developer compares Claude Code, Codex, and Gemini CLI coding agents over a week, noting strengths in context handling, precision, and context size, and weaknesses in cost, ambiguity handling, and consistency.
Should we totally give up on Gemini for coding?
A user reports that Gemini 3.1 Pro significantly underperforms Codex and Claude for coding, likening it to an inexperienced junior developer, and doubts Google's ability to compete in frontier coding models.
@narens: Benchmaxxed
Gemini 3.7 flash outperforms Fable 5, Opus 5, and GPT-5.6 on the Analyst Agent benchmark by Artificial Analysis.
We made Grok 4.5, GPT-5.5, and Claude build the same apps
This article benchmarks Grok 4.5, GPT-5.5, Claude Opus 4.8, and Claude Fable 5 by having each model build three interactive apps (3D Rubik's Cube, particle gravity sandbox, Breakout game) from a single prompt, comparing their one-shot coding capabilities.