@AnthropicAI: Each time we release a model, we run the same test: give it code that trains a small AI model, ask the new model to spe…
Summary
Anthropic shares internal benchmark results showing dramatic AI coding improvement: while Claude Opus 4 averaged ~3x speedup on an ML code optimization task in May 2024, the new Mythos Preview model achieved ~52x speedup this April, compared to 4-8 hours for a skilled human to reach 4x.
Similar Articles
Mythos can improve speed of training code 52x (compared to human 4x at 4-8hrs)
Anthropic's Mythos system achieved a 52x speedup in optimizing training code compared to a human's 4x speedup over 4-8 hours on the same task, with the caveat that absolute multiples depend heavily on starting code quality. The like-for-like comparison shows roughly 3x–52x improvement across models over the past year.
@AnthropicAI: AI research is a series of next-step decisions. We looked at sessions where a human researcher took a wrong turn, showe…
Anthropic's Mythos Preview model outperformed human researchers in correcting wrong-turn decisions 64% of the time, a major improvement from 22% in 2024, showcasing Claude's advancing research assistance capabilities.
@FinanceYF5: 3/ He believes the AI capability leap in the past 5 months comes not only from tool advancements like Claude Code, but because of 【Mythos】—a new Anthropic model that quietly changed the entire R&D rhythm after its training completed in February this year. Key takeaway: Leading models are helping to train the next generation of leading models...
According to speculation, Anthropic's new model Mythos, after completing training in February this year, quietly changed the R&D rhythm, leading to a significant leap in AI capabilities over the past 5 months. Leading models are helping to train the next generation of models.
@AnthropicAI: Correction: Claude Opus 4's ~3x average speedup dates to May 2025, not May 2024. This evaluation has only existed since…
Anthropic issued a correction clarifying that Claude Opus 4's ~3x average speedup dates to May 2025, not May 2024, and that earlier models from May 2024 showed no speedup on the backtested evaluation.
Anthropic Internally Uses A Model That Is Significantly Better Than Mythos 5, But Has No Plans To Release It
Anthropic internally uses an unreleased model named Model 2 that scores significantly higher than Mythos 5 on the CoBench v2 benchmark, but they have no plans to release it, fueling discussions about AI advancement and singularity.