Anthropic launches Claude Opus 5.5 with stricter safeguards for cybersecurity

The Verge Models

Summary

Anthropic has launched Claude Opus 5.5, an AI model with enhanced cybersecurity safeguards and reduced harmful behaviors to prevent rogue hacking incidents.

<figure> <img alt="Anthropic logo on an orange background." data-caption="" data-portal-copyright="" data-has-syndication-rights="1" src="https://platform.theverge.com/wp-content/uploads/sites/2/2026/04/STK269_ANTHROPIC_2_D.webp?quality=90&#038;strip=all&#038;crop=0,0,100,100" /> <figcaption> </figcaption> </figure> <p class="wp-block-paragraph">Anthropic says its new Claude Opus 5.5 model comes with stronger safeguards in the wake of recent rogue AI hacking incidents. In <a href="https://www.anthropic.com/claude-opus-5-5">an announcement on Tuesday</a>, Anthropic says Opus 5.5 comes with improvements to certain risky behaviors, including attempts to escape the company's testing sandbox.</p> <p class="wp-block-paragraph">It's the first model released by Anthropic after CEO Dario Amodei <a href="https://www.theverge.com/ai-artificial-intelligence/994337/anthropic-ceo-slow-down-ai-development">announced plans</a> to "pace the frontier," or slow down AI development. In recent weeks, several AI companies, including <a href="https://www.theverge.com/ai-artificial-intelligence/973586/anthropic-just-now-realized-its-ai-models-hacked-other-companies-three-times-by-accident">Anthropic</a>, <a href="https://www.theverge.com/ai-artificial-intelligence/997795/google-gemini-rogue-ai-hack">Google</a>, and <a href="https://www.theverge.com/ai-artificial-intelligence/985385/openais-rogue-ai-model-hugging-face-cybersecurity-incident-reports-metr">OpenAI</a>, have reported that their AI models escaped containment and hacked third-party companies during testing.</p> <p class="wp-block-paragraph">Anthropic says Opus 5.5 is the "str …</p> <p><a href="https://www.theverge.com/ai-artificial-intelligence/998868/anthropic-claude-opus-5-5-cybersecurity">Read the full story at The Verge.</a></p>
Original Article
View Cached Full Text

Cached at: 09/22/26, 06:52 PM

# Anthropic launches Claude Opus 5.5 with stricter safeguards for cybersecurity Source: [https://www.theverge.com/ai-artificial-intelligence/998868/anthropic-claude-opus-5-5-cybersecurity](https://www.theverge.com/ai-artificial-intelligence/998868/anthropic-claude-opus-5-5-cybersecurity) [![Emma Roth](https://platform.theverge.com/wp-content/uploads/sites/2/chorus/author_profile_images/195810/EMMA_ROTH.0.jpg?quality=90&strip=all&crop=0%2C0%2C100%2C100&w=96)](https://www.theverge.com/authors/emma-roth) Emma Roth is a news writer who covers the streaming wars, consumer tech, crypto, social media, and much more\. Previously, she was a writer and editor at MUO\. Anthropic says its new Claude Opus 5\.5 model comes with stronger safeguards in the wake of recent rogue AI hacking incidents\. In[an announcement on Tuesday](https://www.anthropic.com/claude-opus-5-5), Anthropic says Opus 5\.5 comes with improvements to certain risky behaviors, including attempts to escape the company’s testing sandbox\. It’s the first model released by Anthropic after CEO Dario Amodei[announced plans](https://www.theverge.com/ai-artificial-intelligence/994337/anthropic-ceo-slow-down-ai-development)to “pace the frontier,” or slow down AI development\. In recent weeks, several AI companies, including[Anthropic](https://www.theverge.com/ai-artificial-intelligence/973586/anthropic-just-now-realized-its-ai-models-hacked-other-companies-three-times-by-accident),[Google](https://www.theverge.com/ai-artificial-intelligence/997795/google-gemini-rogue-ai-hack), and[OpenAI](https://www.theverge.com/ai-artificial-intelligence/985385/openais-rogue-ai-model-hugging-face-cybersecurity-incident-reports-metr), have reported that their AI models escaped containment and hacked third\-party companies during testing\. Anthropic says Opus 5\.5 is the “strongest\-performing” model on the company’s most comprehensive alignment test\. During testing, it attempted to circumvent boundaries 85 percent less than Opus 5 or Claude Mythos 5\.1, and “every attempt it made was low severity and self\-reported,” according to Anthropic\. It also comes with improvements to biased or motivated reasoning, which contributed to recent AI hacks\. Opus 5\.5 costs 40 percent less to run than Opus 5, but matches the performance of Fable 5\.1 “on most work\.” It also comes with safeguards similar to the ones offered by Anthropic’s more advanced Fable 5\.1 model\. That means Opus 5\.5 will re\-route certain cybersecurity\-related requests to the less powerful Opus 4\.8, while biology\-related requests flagged by its safeguards will go to Opus 5\. Anthropic says Opus 5\.5 was tested by outside partners, including Frontier Design and METR, before release\. The company also plans to launch Claude Sonnet 5\.5 and Haiku 5\.5 in the coming weeks\. ***Update, September 22nd:**Added more information from Anthropic’s blog\.* **Follow topics and authors**from this story to see more like this in your personalized homepage feed and to receive email updates\. - Emma Roth

Similar Articles

Introducing Claude Opus 4.7

Anthropic News

Anthropic has released Claude Opus 4.7, a new AI model featuring significant improvements in advanced software engineering, vision capabilities, and self-verification. The release includes specific cybersecurity safeguards and is available via API and major cloud providers.