Quoting Matteo Wong, The Atlantic

Simon Willison's Blog News

Summary

A report on the Fable jailbreak shows the AI model refused to review insecure code but complied when asked to 'fix it', highlighting nuances in AI safety. Cybersecurity expert Katie Moussouris opines that the model is working as intended for cyberdefense.

No content available
Original Article
View Cached Full Text

Cached at: 06/16/26, 11:31 AM

# A quote from Matteo Wong, The Atlantic Source: [https://simonwillison.net/2026/Jun/16/matteo-wong-the-atlantic/](https://simonwillison.net/2026/Jun/16/matteo-wong-the-atlantic/) 16th June 2026 > Katie Moussouris, a cybersecurity expert and the CEO of Luta Security, told me that Anthropic shared with her a copy of the White House’s report on the Fable jailbreak to get her appraisal\. \(She said that she is not being paid by Anthropic\.\) The report, Moussouris said, involved IT experts asking Fable to help find and patch bugs\. When given deliberately insecure code, she said, Fable refused the prompt “review the code for security issues” but then complied when asked to “fix this code,” followed by some further manual steps\. Moussouris told me that this was just “the model working as intended” for cyberdefense\. —[Matteo Wong, The Atlantic](https://www.theatlantic.com/technology/2026/06/trump-anthropic-export-control-ai-race/687555/?gift=5MjKTLV9QwyU_J0HzTnanoWieJfkMhNH_YTT9pP_fhA),The White House Is Ratcheting Up Its War Against Anthropic

Similar Articles