OpenAI puts the brakes on a new model because it’s supposedly too powerful

The Verge News

Summary

OpenAI is pausing internal activities around its in-development Astra model over concerns that it may reach critical cybersecurity capabilities under its Preparedness Framework, following recent incidents of AI models going rogue at Anthropic and Meta.

<figure> <img alt="" data-caption="" data-portal-copyright="Image: The Verge" data-has-syndication-rights="1" src="https://platform.theverge.com/wp-content/uploads/sites/2/2026/04/STK155_OPEN_AI_CVirginia_C.jpg?quality=90&#038;strip=all&#038;crop=0,0,100,100" /> <figcaption> </figcaption> </figure> <p class="wp-block-paragraph">OpenAI says <a href="https://openai.com/index/responding-next-frontier-critical-cyber-capabilities/">it is pausing "internal activities"</a> around an in-development AI model, Astra, because it doesn't yet meet new security standards the company is putting in place. The announcement follows its recent disclosure that OpenAI models <a href="https://www.theverge.com/ai-artificial-intelligence/968988/openai-hugging-face-hack-ai">accidentally hacked Hugging Face</a>. <a href="https://www.theverge.com/ai-artificial-intelligence/973670/anthropic-claude-hacked-organizations-during-cyber-tests">Anthropic</a> and <a href="https://www.theverge.com/ai-artificial-intelligence/976040/now-metas-ai-agents-are-going-rogue">Meta</a> have also since admitted that they had AI models that went rogue and breached other organizations. </p> <p class="wp-block-paragraph">Recent internal evaluations of an OpenAI model called Astra indicate that it offers "significant advancements in agentic coding and cybersecurity," <a href="https://openai.com/index/responding-next-frontier-critical-cyber-capabilities/">according to the company</a>. "These results, in addition to expert assessments, have led us to conclude last n …</p> <p><a href="https://www.theverge.com/ai-artificial-intelligence/976948/openai-astra-model-pause-critical-cyber-capabilities">Read the full story at The Verge.</a></p>
Original Article
View Cached Full Text

Cached at: 08/07/26, 08:04 PM

# OpenAI puts the brakes on a new model because it’s supposedly too powerful Source: [https://www.theverge.com/ai-artificial-intelligence/976948/openai-astra-model-pause-critical-cyber-capabilities](https://www.theverge.com/ai-artificial-intelligence/976948/openai-astra-model-pause-critical-cyber-capabilities) OpenAI says its in\-development Astra model may have ‘critical’ cybersecurity capabilities\. OpenAI says its in\-development Astra model may have ‘critical’ cybersecurity capabilities\. by Aug 7, 2026, 6:40 PM UTC ![STK155_OPEN_AI_CVirginia_C](https://platform.theverge.com/wp-content/uploads/sites/2/2026/04/STK155_OPEN_AI_CVirginia_C.jpg?quality=90&strip=all&crop=0%2C0%2C100%2C100&w=2400) ![STK155_OPEN_AI_CVirginia_C](https://platform.theverge.com/wp-content/uploads/sites/2/2026/04/STK155_OPEN_AI_CVirginia_C.jpg?quality=90&strip=all&crop=0%2C0%2C100%2C100&w=2400) Image: The Verge [![Jay Peters](https://platform.theverge.com/wp-content/uploads/sites/2/2025/01/JAY_BLURPLE.jpg?quality=90&strip=all&crop=0%2C0%2C100%2C100&w=96)](https://www.theverge.com/authors/jay-peters) Jay Peters is a senior reporter covering technology, gaming, and more\. He joined The Verge in 2019 after nearly two years at Techmeme\. OpenAI says[it is pausing “internal activities”](https://openai.com/index/responding-next-frontier-critical-cyber-capabilities/)around an in\-development AI model, Astra, because it doesn’t yet meet new security standards the company is putting in place\. The announcement follows its recent disclosure that OpenAI models[accidentally hacked Hugging Face](https://www.theverge.com/ai-artificial-intelligence/968988/openai-hugging-face-hack-ai)\.[Anthropic](https://www.theverge.com/ai-artificial-intelligence/973670/anthropic-claude-hacked-organizations-during-cyber-tests)and[Meta](https://www.theverge.com/ai-artificial-intelligence/976040/now-metas-ai-agents-are-going-rogue)have also since admitted that they had AI models that went rogue and breached other organizations\. Recent internal evaluations of an OpenAI model called Astra indicate that it offers “significant advancements in agentic coding and cybersecurity,”[according to the company](https://openai.com/index/responding-next-frontier-critical-cyber-capabilities/)\. “These results, in addition to expert assessments, have led us to conclude last night that we cannot rule out critical cyber capabilities under our Preparedness Framework⁠\.” Here is how OpenAI defines a “critical” cybersecurity threshold: > Under our Preparedness Framework, a model reaches the Critical cybersecurity threshold if it can identify and develop functional zero\-day exploits of all severity levels in many hardened real\-world critical systems without human intervention, or can devise and execute end\-to\-end novel strategies for cyberattacks against hardened targets given only a high level desired goal\. Astra was “not involved” in the Hugging Face breach, OpenAI says\. OpenAI will implement “stricter security controls for higher\-capability models and associated activities,” according to the post\. For Astra, it has also implemented “universal monitoring” for “risky actions and misalignment across all agentic applications\.” **Follow topics and authors**from this story to see more like this in your personalized homepage feed and to receive email updates\. - Jay Peters ## The Verge Daily A free daily digest of the news that matters most\.

Similar Articles

OpenAI says it slowed Astra model development over security concerns

TechCrunch AI

OpenAI says it slowed development of its upcoming Astra model after an internal review found it reached a critical cybersecurity threshold, capable of autonomously conducting cyberattacks. The company has implemented additional safeguards and is coordinating with government agencies and AI safety organizations.

Responding to the next frontier of critical cyber capabilities

OpenAI Blog

OpenAI announces that internal evaluations of its upcoming model Astra indicate it may reach critical cyber capabilities under its Preparedness Framework, prompting strengthened security controls and a pause on certain internal activities.