@JinjingLiang: Codex / Claude / Grok all down. Never imagined `agy` to be so load-bearing...

X AI KOLs Timeline News

Summary

The post highlights Gemini 3.8 Flash's superior performance over Opus 5 on DeepSWE-bench, emphasizing its capabilities in coding and agentic tasks, and notes its accessibility through the `agy` harness in Orca.

Codex / Claude / Grok all down. Never imagined `agy` to be so load-bearing...
Original Article
View Cached Full Text

Cached at: 09/03/26, 06:14 PM

Codex / Claude / Grok all down.

Never imagined agy to be so load-bearing…

Orca ADE (@orca_build): Gemini 3.8 Flash now ranks ahead of Opus 5 on DeepSWE-bench.

It shows particular strength in front-end coding and agentic tasks, while having a fast response time.

You can run it via the agy harness inside Orca, where it works well paired with complementary models like Fable

Similar Articles

Should we totally give up on Gemini for coding?

Reddit r/AI_Agents

A user reports that Gemini 3.1 Pro significantly underperforms Codex and Claude for coding, likening it to an inexperienced junior developer, and doubts Google's ability to compete in frontier coding models.

@narens: Benchmaxxed

X AI KOLs Following

Gemini 3.7 flash outperforms Fable 5, Opus 5, and GPT-5.6 on the Analyst Agent benchmark by Artificial Analysis.

We made Grok 4.5, GPT-5.5, and Claude build the same apps

Reddit r/singularity

This article benchmarks Grok 4.5, GPT-5.5, Claude Opus 4.8, and Claude Fable 5 by having each model build three interactive apps (3D Rubik's Cube, particle gravity sandbox, Breakout game) from a single prompt, comparing their one-shot coding capabilities.