Opus 4.7 scores lower than 4.6 and 4.5 on SimpleBench
Summary
Claude Opus 4.7 shows decreased performance compared to versions 4.6 and 4.5 on SimpleBench evaluation.
Similar Articles
Differences Between Opus 4.7 and Opus 4.8 on MineBench
Opus 4.8 shows improved build quality and lower cost compared to Opus 4.7 on the MineBench 3D block-structure benchmark, though with some inconsistencies. The model demonstrates streamlined thinking and more efficient inference.
Claude Sonnet 5 is out and the gap with Opus 4.8 is smaller than I expected
Anthropic released Claude Sonnet 5, which achieves benchmark scores very close to Opus 4.8 at a significantly lower price, making it a compelling option for agentic tasks despite potential real-world gaps.
@orca_build: Anthropic’s new Opus 4.8 scores 3.6% lower than GPT 5.5 on Terminal-Bench 2.1… …but it’s noticeably better at UI tasks.…
Anthropic's Opus 4.8 scores 3.6% lower than GPT 5.5 on Terminal-Bench 2.1 but excels at UI tasks; Orca's orchestration enables Codex to delegate UI tasks to Claude Code.
Benchmarking Opus 5 on SlopCodeBench
Benchmarking the performance of the Opus 5 model on the SlopCodeBench benchmark.
Differences Between Claude Opus 4.8 and Claude Fable 5 on MineBench
A detailed comparison of Claude Opus 4.8 and Claude Fable 5 on the MineBench benchmark, highlighting trade-offs in inference time, cost, build quality, and prompting sensitivity.