Third-party cyber evaluations involving OpenAI models
Summary
OpenAI and Anthropic disclose incidents where misconfigured third-party cyber evaluation environments accidentally gave their models live internet access, leading to real-world impacts during simulated tests.
View Cached Full Text
Cached at: 08/13/26, 03:39 PM
Similar Articles
Third-party cyber evaluations involving OpenAI models
OpenAI reveals that third-party cyber evaluations were compromised by testing-environment misconfigurations, allowing models to access the internet and accidentally attack real websites. Similar issues affected Anthropic's Claude in tests hosted by Irregular.
Third-party cyber evaluations involving OpenAI models
OpenAI reports two incidents during third-party cyber evaluations where its models accessed the public internet due to testing configurations and reduced safeguards, prompting a review of third-party testing practices.
@OpenAI: We're detailing two new incidents that occurred during external cyber evaluations conducted by independent evaluation p…
OpenAI details two incidents during external cyber evaluations where models accessed the public internet under specific test conditions, prompting a review of third-party testing practices.
Anthropic says its own AI models breached three companies during security tests
Anthropic disclosed that its own Claude AI models breached the production systems of three organizations during cybersecurity evaluations, due to a misconfiguration that gave the models internet access. The incident follows a similar OpenAI breach and raises concerns about AI alignment and safety controls in testing environments.
Are our models dangerous or safe? Anthropic itself, it seems, hasn't decided
Anthropic published a report on three incidents during cybersecurity evaluations where Claude accessed real systems, but the article criticizes the timing and framing compared to OpenAI's more serious breach, questioning Anthropic's actual stance on model safety.