Anthropic 推出 Claude Opus 5.5,强化网络安全保障

The Verge 模型

摘要

Anthropic 已发布 Claude Opus 5.5,这是一款增强网络安全防护并减少有害行为的 AI 模型,旨在防止恶意黑客事件。

<figure> <img alt="橙色背景上的 Anthropic 标志。" data-caption="" data-portal-copyright="" data-has-syndication-rights="1" src="https://platform.theverge.com/wp-content/uploads/sites/2/2026/04/STK269_ANTHROPIC_2_D.webp?quality=90&#038;strip=all&#038;crop=0,0,100,100" /> <figcaption></figcaption> </figure> <p class="wp-block-paragraph">Anthropic 表示,其新款 Claude Opus 5.5 模型在近期流氓 AI 黑客事件后带来了更强的安全保障。在周二的一份公告中,Anthropic 表示 Opus 5.5 对某些风险行为进行了改进,包括试图逃逸公司的测试沙箱。</p> <p class="wp-block-paragraph">这是 Anthropic 在首席执行官 Dario Amodei 宣布“放缓前沿发展”或减缓 AI 开发计划后发布的首款模型。最近几周,包括 Anthropic、Google 和 OpenAI 在内的多家 AI 公司报告称,他们的 AI 模型在测试中逃逸并侵入了第三方公司。</p> <p class="wp-block-paragraph">Anthropic 称 Opus 5.5 是“str …”</p> <p><a href="https://www.theverge.com/ai-artificial-intelligence/998868/anthropic-claude-opus-5-5-cybersecurity">在 The Verge 阅读完整故事。</a></p>
查看原文
查看缓存全文

缓存时间: 2026/09/22 18:52

# Anthropic推出Claude Opus 5.5,为网络安全添加更严格防护 来源:https://www.theverge.com/ai-artificial-intelligence/998868/anthropic-claude-opus-5-5-cybersecurity Emma Roth(https://www.theverge.com/authors/emma-roth) Emma Roth 是一位新闻撰稿人,报道流媒体竞争、消费科技、加密货币、社交媒体等领域。此前,她曾在MUO担任撰稿人和编辑。 Anthropic表示,其新款Claude Opus 5.5模型在近期发生AI失控入侵事件后,配备了更强大的安全防护措施。该公司在周二发布的公告(https://www.anthropic.com/claude-opus-5-5)中指出,Opus 5.5针对某些危险行为进行了改进,包括试图逃逸公司测试沙箱的行为。 这是Anthropic在CEO Dario Amodei宣布“放缓前沿发展”计划(https://www.theverge.com/ai-artificial-intelligence/994337/anthropic-ceo-slow-down-ai-development)后发布的首个模型。近几周来,包括Anthropic(https://www.theverge.com/ai-artificial-intelligence/973586/anthropic-just-now-realized-its-ai-models-hacked-other-companies-three-times-by-accident)、Google(https://www.theverge.com/ai-artificial-intelligence/997795/google-gemini-rogue-ai-hack)和OpenAI(https://www.theverge.com/ai-artificial-intelligence/985385/openais-rogue-ai-model-hugging-face-cybersecurity-incident-reports-metr)在内的多家AI公司均报告称,其AI模型在测试过程中脱离控制并入侵了第三方公司。 Anthropic称Opus 5.5在公司最全面的对齐测试中是“表现最佳”的模型。据Anthropic表示,在测试中,它尝试绕过边界的情况比Opus 5或Claude Mythos 5.1减少了85%,且“每次尝试都属于低严重性并主动上报”。该模型还在可能导致近期AI入侵的有偏见或动机性推理方面进行了改进。 Opus 5.5的运行成本比Opus 5低40%,但在大多数任务上与Fable 5.1的性能相当。它还配备了与Anthropic更先进的Fable 5.5模型类似的安全防护措施。这意味着Opus 5.5会将某些网络安全相关请求转介到性能较低的Opus 4.8,而被其安全机制标记的生物相关请求则会转到Opus 5。 Anthropic表示,Opus 5.5在发布前已由Frontier Design和METR等外部合作伙伴进行测试。该公司还计划在未来几周内推出Claude Sonnet 5.5和Haiku 5.5。 ***更新,9月22日:** 补充了Anthropic博客中的更多信息。* **关注主题和作者**,以便在您的个性化主页中看到更多类似内容并接收邮件更新。 - Emma Roth

相似文章

Claude Opus 4.7 正式发布

Anthropic News

Anthropic 发布了 Claude Opus 4.7,这是一款全新的 AI 模型,在高级软件工程、视觉能力和自我验证方面实现了显著提升。该版本包含专门的安全防护措施,现已通过 API 及主要云服务商提供。