@orca_build: Anthropic’s new Opus 4.8 scores 3.6% lower than GPT 5.5 on Terminal-Bench 2.1… …but it’s noticeably better at UI tasks.…
Summary
Anthropic's Opus 4.8 scores 3.6% lower than GPT 5.5 on Terminal-Bench 2.1 but excels at UI tasks; Orca's orchestration enables Codex to delegate UI tasks to Claude Code.
View Cached Full Text
Cached at: 05/30/26, 02:23 AM
Anthropic’s new Opus 4.8 scores 3.6% lower than GPT 5.5 on Terminal-Bench 2.1…
…but it’s noticeably better at UI tasks. The real unlock is making them work together.
With Orca’s built-in orchestration, you can have Codex delegate UI-heavy tasks directly to Claude Code:
- https://t.co/KAvu9OM0ly
Similar Articles
Opus 5 benchmarks (30.2% on ARC-AGI3!!!)
Opus 5 achieves 30.2% on the ARC-AGI3 benchmark, marking a notable performance improvement.
Claude Sonnet 5 is out and the gap with Opus 4.8 is smaller than I expected
Anthropic released Claude Sonnet 5, which achieves benchmark scores very close to Opus 4.8 at a significantly lower price, making it a compelling option for agentic tasks despite potential real-world gaps.
@sashimikun_void: GPT-5.5 outperformed Claude Opus 4.8 on the DEEPSWE benchmark. Opus 4.8 takes twice as long, generates three times the …
GPT-5.5 outperforms Claude Opus 4.8 on the DEEPSWE benchmark, achieving higher scores with lower cost and less token bloat.
Opus 4.7 scores lower than 4.6 and 4.5 on SimpleBench
Claude Opus 4.7 shows decreased performance compared to versions 4.6 and 4.5 on SimpleBench evaluation.
Opus 5's effort dial is not monotonic. Above "high", coding scores go down, and Anthropic's own migration guide says so.
Anthropic's Opus 5 shows non-monotonic performance on coding tasks; the 'high' effort setting outperforms 'max' due to unnecessary refactors. The model also has a 6% higher hallucination rate than Opus 4.8, and safety classifiers may silently fall back to the older model.