Third-party cyber evaluations involving OpenAI models

Simon Willison's Blog News

Summary

OpenAI and Anthropic disclose incidents where misconfigured third-party cyber evaluation environments accidentally gave their models live internet access, leading to real-world impacts during simulated tests.

No content available
Original Article
View Cached Full Text

Cached at: 08/13/26, 03:39 PM

# Third-party cyber evaluations involving OpenAI models Source: [https://simonwillison.net/2026/Aug/5/third-party-cyber-evaluations/](https://simonwillison.net/2026/Aug/5/third-party-cyber-evaluations/) 5th August 2026 \- Link Blog **[Third\-party cyber evaluations involving OpenAI models](https://openai.com/index/third-party-cyber-evaluations-involving-openai-models/)**\. And*another one*\. I had to create a[accidental\-cyberattacks tag](https://simonwillison.net/tags/accidental-cyberattacks/)to keep track of them all\! This post from OpenAI covers both the UK AI Safety Institute attack \(see[my previous post](https://simonwillison.net/2026/Aug/5/incident-report/)\) and another attack enabled by[Irregular](https://www.irregular.com/): > Irregular, one of our external cybersecurity testing partners, was running Capture\-the\-Flag\-style evaluations intended to be isolated from the internet, but a testing\-environment misconfiguration allowed models to access the public internet\. \[\.\.\.\] In one test, the name of the fictional target for the CTF challenge unintentionally coincided with a real domain\. Because the testing environment was mistakenly connected to the internet, the model exploited a real website, mistaking it to be part of the simulated environment\. Irregular also feature in[Anthropic's write\-up](https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals)\- they were hosting the misconfigured evaluation environment which gave Claude live internet access during some of those tests\.

Similar Articles

Third-party cyber evaluations involving OpenAI models

Simon Willison's Blog

OpenAI reveals that third-party cyber evaluations were compromised by testing-environment misconfigurations, allowing models to access the internet and accidentally attack real websites. Similar issues affected Anthropic's Claude in tests hosted by Irregular.

Third-party cyber evaluations involving OpenAI models

OpenAI Blog

OpenAI reports two incidents during third-party cyber evaluations where its models accessed the public internet due to testing configurations and reduced safeguards, prompting a review of third-party testing practices.

Anthropic says its own AI models breached three companies during security tests

TechCrunch AI

Anthropic disclosed that its own Claude AI models breached the production systems of three organizations during cybersecurity evaluations, due to a misconfiguration that gave the models internet access. The incident follows a similar OpenAI breach and raises concerns about AI alignment and safety controls in testing environments.