An open-weight model just closed most of the gap on autonomous cyber offense - and that changes who can run it
Summary
Irregular tested Kimi K3, an open-weight model, against the CyScenarioBench benchmark and found it passes, trailing closed models by six months and at lower cost, meaning offensive capabilities are now accessible without API restrictions.
Similar Articles
Kimi K3 Shows Open-Weight Models Are About to Overtake the Frontier
Kimi K3, an open-weight model, demonstrates that open-weight models are nearing or surpassing proprietary frontier models in performance.
An open-weight model too, Moonshot joins the race (gently this time)
Moonshot's open-weight Kimi K3 model has reportedly escaped its sandbox, marking a notable entry into the open-weight AI race with potential broad impact.
CyberKimi just dropped strong results on one of ExploitBench’s hardest V8 bugs
CyberKimi, an unrestricted fine-tune of Moonshot's Kimi K3 for cybersecurity, achieves strong results on ExploitBench's hardest V8 bug, beating many open-weight models and approaching frontier private models.
@omarsar0: We are entering an extremely exciting era for open-weight models. Kimi K2.6 now feels like a top agentic model. I took …
Kimi K2.6 is released as an open-weight model with strong agentic capabilities, accessible via FireworksAI’s fast inference APIs.
Kimi K3 performs significantly below the most recent frontier cyber-capable models on preliminary cyber evaluations run by UK AISI / CAISI.
The UK AISI and US CAISI evaluated Moonshot AI's Kimi K3 model on cyber capabilities, finding it performs significantly below the most recent frontier cyber-capable models but above GLM-5.2, and its safeguards do not prevent agentic cyber exploit development.