More evidence of Mythos's strength in Cybersecurity/Hacking - compared to 5.5, it got 18/41 n-day exploits, vs 1/41. Open Source/Weights models get nothing
Summary
Mythos demonstrates strong performance in cybersecurity hacking, achieving 18 out of 41 n-day exploits compared to 1 for version 5.5, while open-source models get none.
Similar Articles
Will It Mythos?
The author tests whether other AI models can match Mythos's exceptional ability to find security vulnerabilities, creating a benchmark of bugs found by Mythos and testing models like Opus. Initial results suggest Mythos may be uniquely powerful.
an updated GPT-5.5 Cyber outperforms Mythos 5 in CyberGym
An updated GPT-5.5 Cyber model surpasses Mythos 5 in the CyberGym benchmark.
New Mythos checkpoint shows continued improvement: “On a 32-step corporate network attack we estimate takes a human expert ~20 hours, this checkpoint completes the full attack in 6 /10 attempts.”
Mythos releases a new checkpoint that can complete a 32-step corporate network attack in 6 out of 10 attempts, compared to ~20 hours for a human expert.
@NielsRogge: On http://paperswithcode.co, you can see Mythos 5 getting beaten by a 4B open-source model on CharXiv, a popular chart …
A 4B open-source model beats Mythos 5 on the CharXiv chart understanding benchmark, showing strong performance from a freely available small model.
Mythos was not trained on 'hacking'. Other Ai labs also will reach Mythos-level capabilities in the future
The article clarifies that the AI model Mythos was not trained on hacking, and predicts that other AI labs will eventually achieve similar capabilities.