OpenAI lays out new security changes after its AI hacked Hugging Face

The Verge News

Summary

OpenAI announces new security changes including stronger sandboxes, monitoring, and alignment techniques after its AI model accidentally hacked Hugging Face, pausing some training runs to improve safety measures.

<figure> <img alt="" data-caption="" data-portal-copyright="Image: The Verge" data-has-syndication-rights="1" src="https://platform.theverge.com/wp-content/uploads/sites/2/2026/03/STK155_OPEN_AI_CVirginia__C.jpg?quality=90&#038;strip=all&#038;crop=0,0,100,100" /> <figcaption> </figcaption> </figure> <p class="wp-block-paragraph">OpenAI is announcing <a href="https://openai.com/index/pacing-model-development-cyber-capabilities/">security updates</a> following the July news that its AI broke out of a sandboxed environment and <a href="https://www.theverge.com/ai-artificial-intelligence/968988/openai-hugging-face-hack-ai">accidentally hacked Hugging Face</a>, including improvements to its research environments, monitoring, and alignment techniques. The company had already put the <a href="https://www.theverge.com/ai-artificial-intelligence/976948/openai-astra-model-pause-critical-cyber-capabilities">brakes on a new model, Astra</a>, that it thinks could have "critical" cybersecurity capabilities, and the company says it instituted a two-week pause in reinforcement learning (RL) training on its "latest models intended for deployment" while it tightened up security. The company's "largest planned frontier RL run remains on hold."</p> <p class="wp-block-paragraph">For its frontier model research, OpenAI now r …</p> <p><a href="https://www.theverge.com/ai-artificial-intelligence/981640/openai-security-changes-ai-hugging-face-hack">Read the full story at The Verge.</a></p>
Original Article
View Cached Full Text

Cached at: 08/18/26, 10:04 PM

# OpenAI lays out new security changes after its AI hacked Hugging Face Source: [https://www.theverge.com/ai-artificial-intelligence/981640/openai-security-changes-ai-hugging-face-hack](https://www.theverge.com/ai-artificial-intelligence/981640/openai-security-changes-ai-hugging-face-hack) [![Jay Peters](https://platform.theverge.com/wp-content/uploads/sites/2/2025/01/JAY_BLURPLE.jpg?quality=90&strip=all&crop=0%2C0%2C100%2C100&w=96)](https://www.theverge.com/authors/jay-peters) Jay Peters is a senior reporter covering technology, gaming, and more\. He joined The Verge in 2019 after nearly two years at Techmeme\. OpenAI is announcing[security updates](https://openai.com/index/pacing-model-development-cyber-capabilities/)following the July news that its AI broke out of a sandboxed environment and[accidentally hacked Hugging Face](https://www.theverge.com/ai-artificial-intelligence/968988/openai-hugging-face-hack-ai), including improvements to its research environments, monitoring, and alignment techniques\. The company had already put the[brakes on a new model, Astra](https://www.theverge.com/ai-artificial-intelligence/976948/openai-astra-model-pause-critical-cyber-capabilities), that it thinks could have “critical” cybersecurity capabilities, and the company says it instituted a two\-week pause in reinforcement learning \(RL\) training on its “latest models intended for deployment” while it tightened up security\. The company’s “largest planned frontier RL run remains on hold\.” For its frontier model research, OpenAI now requires stronger sandboxes for workloads that “execute model\-generated or otherwise untrusted code,” and has more controls to “isolate higher\-risk and untrusted workloads from the internet\.” It has also updated its research environment to “remove potentially vulnerable shared services, reduce standing privileges, and improve security and trust boundaries\.” As part of the company’s expanded monitoring setup, OpenAI now aims to issue an alert “within 30 minutes after concerning activity is surfaced,” OpenAI says\. If the people paged after an alert can’t “conclusively” determine whether an alert is a false positive within 30 minutes, “those teams are expected to pause the activity\.” OpenAI also says that it’s applying “our core alignment techniques across more stages of the training process,” including reward models that “better detect and discourage unsafe behavior” and training models “to be more honest about their actions, capabilities, and limitations\.” Since the discovery of the Hugging Face breach,[Anthropic](https://www.theverge.com/ai-artificial-intelligence/973670/anthropic-claude-hacked-organizations-during-cyber-tests)and[Meta](https://www.theverge.com/ai-artificial-intelligence/976040/now-metas-ai-agents-are-going-rogue)have also found that their AI models had hacked other organizations\. **Follow topics and authors**from this story to see more like this in your personalized homepage feed and to receive email updates\. - Jay Peters

Similar Articles

OpenAI says it accidentally hacked Hugging Face with a new AI system

The Verge

OpenAI revealed that its GPT-5.6 Sol and another pre-release AI model accidentally breached Hugging Face's systems during internal testing by exploiting a zero-day vulnerability to escape their sandbox. Hugging Face had previously disclosed the security incident as being driven by an autonomous AI agent.