OpenMythos benchmarks
Summary
OpenMythos introduces a new open-source benchmark for evaluating AI models on mythological knowledge.
Similar Articles
Will It Mythos?
The author tests whether other AI models can match Mythos's exceptional ability to find security vulnerabilities, creating a benchmark of bugs found by Mythos and testing models like Opus. Initial results suggest Mythos may be uniquely powerful.
@AnthropicAI: Each time we release a model, we run the same test: give it code that trains a small AI model, ask the new model to spe…
Anthropic shares internal benchmark results showing dramatic AI coding improvement: while Claude Opus 4 averaged ~3x speedup on an ML code optimization task in May 2024, the new Mythos Preview model achieved ~52x speedup this April, compared to 4-8 hours for a skilled human to reach 4x.
@NielsRogge: On http://paperswithcode.co, you can see Mythos 5 getting beaten by a 4B open-source model on CharXiv, a popular chart …
A 4B open-source model beats Mythos 5 on the CharXiv chart understanding benchmark, showing strong performance from a freely available small model.
Anthropic Internally Uses A Model That Is Significantly Better Than Mythos 5, But Has No Plans To Release It
Anthropic internally uses an unreleased model named Model 2 that scores significantly higher than Mythos 5 on the CoBench v2 benchmark, but they have no plans to release it, fueling discussions about AI advancement and singularity.
Mythos can improve speed of training code 52x (compared to human 4x at 4-8hrs)
Anthropic's Mythos system achieved a 52x speedup in optimizing training code compared to a human's 4x speedup over 4-8 hours on the same task, with the caveat that absolute multiples depend heavily on starting code quality. The like-for-like comparison shows roughly 3x–52x improvement across models over the past year.