OpenMythos benchmarks

Reddit r/LocalLLaMA Papers

Summary

OpenMythos introduces a new open-source benchmark for evaluating AI models on mythological knowledge.

No content available
Original Article

Similar Articles

Will It Mythos?

Hacker News Top

The author tests whether other AI models can match Mythos's exceptional ability to find security vulnerabilities, creating a benchmark of bugs found by Mythos and testing models like Opus. Initial results suggest Mythos may be uniquely powerful.

Mythos can improve speed of training code 52x (compared to human 4x at 4-8hrs)

Reddit r/singularity

Anthropic's Mythos system achieved a 52x speedup in optimizing training code compared to a human's 4x speedup over 4-8 hours on the same task, with the caveat that absolute multiples depend heavily on starting code quality. The like-for-like comparison shows roughly 3x–52x improvement across models over the past year.