@jerryjliu0: An interesting thing I'm observing from the blog/system card is that on a good chunk of the reported benchmarks (~20-30…

X AI KOLs Following Models

Summary

An observation that Claude Opus 5's max thinking leads to performance degradation on ~20-30% of benchmarks compared to xhigh, contrary to the expectation that more test-time compute improves performance.

An interesting thing I'm observing from the blog/system card is that on a good chunk of the reported benchmarks (~20-30% from a skim), Opus 5 max thinking leads to a degradation in performance compared to xhigh. Usually you would assume that as you increase thinking and test-time compute, performance goes up. Some of these results contradict that assumption. I wonder if this is a natural emergent property of smaller models or a posttraining issue.
Original Article
View Cached Full Text

Cached at: 07/25/26, 08:03 AM

An interesting thing I’m observing from the blog/system card is that on a good chunk of the reported benchmarks (~20-30% from a skim), Opus 5 max thinking leads to a degradation in performance compared to xhigh.

Usually you would assume that as you increase thinking and test-time compute, performance goes up. Some of these results contradict that assumption. I wonder if this is a natural emergent property of smaller models or a posttraining issue.

Claude (@claudeai): Introducing Claude Opus 5.

It’s a thoughtful and proactive model that comes close to the frontier intelligence of Fable 5 at half the price.

Similar Articles

@polynoamial: https://x.com/polynoamial/status/2064210146558136827

X AI KOLs Following

This article argues that LLM benchmark performance is increasingly a function of test-time compute, and that current evaluation methods fail to capture capability improvements when controlling for inference budget. It advocates for plotting performance vs. tokens, cost, or time, and discusses implications for safety evaluations.

@injaneity: https://x.com/injaneity/status/2075659478096376158

X AI KOLs Timeline

This article explains how batching and parallel operations improve latency and efficiency in AI computer use systems, highlighting open-source implementations like pi-computer-use and cua-driver that achieved significant performance gains before similar features appeared in Codex.