Opus 5 Great Performance -> Gaslighting
Summary
A user criticizes Opus 5, alleging that its performance is poor compared to claims and that positive reviews are from non-serious testers or bots.
Similar Articles
@robinebers: Fable 5 Low > Opus 4.8 Max
A user posts a comparison suggesting Fable 5 Low outperforms Opus 4.8 Max, with another user commenting that Fable 5 is back but being used incorrectly.
Opus 5 allegedly produced these outputs
Rumored outputs from an AI model called Opus 5 have surfaced, but their authenticity is unconfirmed.
@orca_build: Anthropic’s new Opus 4.8 scores 3.6% lower than GPT 5.5 on Terminal-Bench 2.1… …but it’s noticeably better at UI tasks.…
Anthropic's Opus 4.8 scores 3.6% lower than GPT 5.5 on Terminal-Bench 2.1 but excels at UI tasks; Orca's orchestration enables Codex to delegate UI tasks to Claude Code.
Opus 5 benchmarks (30.2% on ARC-AGI3!!!)
Opus 5 achieves 30.2% on the ARC-AGI3 benchmark, marking a notable performance improvement.
Opus 5's effort dial is not monotonic. Above "high", coding scores go down, and Anthropic's own migration guide says so.
Anthropic's Opus 5 shows non-monotonic performance on coding tasks; the 'high' effort setting outperforms 'max' due to unnecessary refactors. The model also has a 6% higher hallucination rate than Opus 4.8, and safety classifiers may silently fall back to the older model.