@FradSer: The most interesting thing I've done so far: Trying a series of methods to make models like gpt-oss:20b and gemma4:e4b approach Opus 4.7's level under certain conditions

X AI KOLs Timeline News

Summary

Attempting a series of methods to make models such as gpt-oss:20b and gemma4:e4b approach Opus 4.7's performance level under certain conditions.

The most interesting thing I've done so far: Trying a series of methods to make models like gpt-oss:20b and gemma4:e4b approach Opus 4.7's level under certain conditions 👀 https://t.co/1YUmoZ8dao
Original Article
View Cached Full Text

Cached at: 05/24/26, 02:18 AM

The most interesting thing I’ve done so far:

Trying a series of approaches to get models like gpt-oss:20b and gemma4:e4b close to Opus 4.7 level under certain conditions 👀 https://t.co/1YUmoZ8dao

Similar Articles

@FuckAnthropic: Conducted a comparative analysis. Overall, DeepSeek V4 Flash-0731 is roughly a model at the level between Opus 4.7 and 4.8, entering the frontier Agent model competition with a minimal activation scale, and at about 1/12 to 1/60 of the token cost to enter the frontier Ag…

X AI KOLs Timeline

The author's comparative analysis concludes that DeepSeek V4 Flash-0731 achieves Opus 4.7–4.8 level performance with an extremely small activation scale, entering the frontier agent model tier at a very low token cost. It surpasses GLM-5.2 overall, but its shortfalls remain difficult repository-level coding and long-horizon engineering.

Gemma 4 31B's competence surprised me

Reddit r/LocalLLaMA

A user shares anecdotal findings that Gemma 4 31B outperforms Qwen 3.6 models and matches Opus 4.7 in understanding and refactoring messy academic code, highlighting a benchmark (SciCode) where Gemma excels.

@RealYDT: Holy shit, brothers, I suddenly have a bit of a conspiracy theory, but it makes more sense the more I think about it: Maybe Opus 4.6 was the last generation of Opus that was actually used extensively by human employees internally at Anthropic. After that, the internal focus shifted to Mythos. Then, the training and RL... for 4.7, 4.8, and 5 were increasingly done by a stronger "AI teacher" like Mythos to spot errors, score, and correct.

X AI KOLs Following

The user speculates that starting from Opus 4.6, Anthropic internally shifted to using the Mythos model, leading to subsequent Opus models being trained and behaving more cautiously and AI-like.