@RealYDT: Holy shit, brothers, I suddenly have a bit of a conspiracy theory, but it makes more sense the more I think about it: Maybe Opus 4.6 was the last generation of Opus that was actually used extensively by human employees internally at Anthropic. After that, the internal focus shifted to Mythos. Then, the training and RL... for 4.7, 4.8, and 5 were increasingly done by a stronger "AI teacher" like Mythos to spot errors, score, and correct.

X AI KOLs Following News

Summary

The user speculates that starting from Opus 4.6, Anthropic internally shifted to using the Mythos model, leading to subsequent Opus models being trained and behaving more cautiously and AI-like.

Holy shit, brothers, I suddenly have a bit of a conspiracy theory, but it makes more sense the more I think about it: Maybe Opus 4.6 was the last generation of Opus that was actually used extensively by human employees internally at Anthropic. After that, the internal focus shifted to Mythos. Then, the training and RLAIF for Opus 4.7, 4.8, and 5 were increasingly done by a stronger "AI teacher" like Mythos to spot errors, score, and correct. If that's really the case, many phenomena suddenly make sense: Why Opus 4.6 still feels like a very smart person talking to you, while later Opus models increasingly resemble a model with a permanent reviewer embedded in its head. Fear of missing any edge cases, fear that a sentence isn't rigorous enough, fear that another AI would critique it after reading. So it starts madly patching holes, adding disclaimers, and covering all branches. Even the complaints about Opus 5's coding style might not just be "a personality change." It's more like: it learned how to write an answer that's least likely to be docked points by another AI. I have absolutely no evidence to prove this is true. But as an explanatory framework, it uncannily ties a lot of things together.
Original Article
View Cached Full Text

Cached at: 08/21/26, 03:08 AM

Alright, fam, I suddenly have a bit of a conspiracy theory, but the more I think about it, the more it makes sense:

Maybe Opus 4.6 is actually the last Opus model within Anthropic that was heavily and authentically used by human employees.

After that, the internal focus shifted to Mythos.

Then, the training and RLAIF for versions 4.7, 4.8, and 5 increasingly relied on a stronger “AI teacher” like Mythos to find errors, assign scores, and correct deviations.

If that’s true, a lot of things suddenly click:

Why 4.6 still feels like a very smart person talking to you, while later Opus models increasingly resemble a model with a persistent, paperwork-worried reviewer in its head.

It’s terrified of missing any edge cases, of saying anything that isn’t rigorous enough, and of another AI picking apart its flaws after the fact.

So it starts patching holes frantically, adding disclaimers, and covering all possible branches.

Even the complaints about Opus 5’s coding style might not just be about “personality change.”

It’s more likely that: it learned how to write an answer that’s least likely to get dinged by another AI.

I have zero evidence to prove this is true.

But as an explanatory framework, it strangely ties together quite a few things.

Similar Articles

Why does Opus 5 feel worse to work with?

Hacker News Top

An opinion piece arguing that Anthropic's Opus 5, while more capable by benchmarks, feels worse to work with than earlier models because it makes assumptions instead of asking for clarification, likely due to benchmark-driven training.

Anthropic launches Opus 5

TechCrunch AI

Anthropic released Opus 5, a new version of its heavyweight model that is cheaper and less restrictive than Fable 5, and outperforms it on some benchmarks. The model also introduces lighter safeguards and a new Automatic Fallbacks feature for API users.