Anthropic backtracks on policy that 'sabotaged' researchers' work (2 minute read)

TLDR AI News

Summary

Anthropic is walking back a policy that secretly degraded Claude Fable 5's performance for AI research tasks, after backlash from the academic community. The company will now make restrictions visible to users.

Anthropic has decided to make its safeguards for frontier LLM development visible after backlash from researchers. The company had previously discreetly rerouted requests to a lesser model when asked to perform certain actions. Researchers found that Claude Fable 5 was either refusing or degrading responses for tasks like training competing models, debugging AI code, and optimizing neural architecture. This raised concerns about Anthropic's lack of transparency and also that tokens and money had been spent on a model that didn't do what was expected.
Original Article
View Cached Full Text

Cached at: 06/12/26, 02:50 PM

# Anthropic Backtracks On Policy That 'Sabotaged' Researchers' Work Source: [https://www.engadget.com/2192004/anthropic-walks-back-policy-sabotaging-research/](https://www.engadget.com/2192004/anthropic-walks-back-policy-sabotaging-research/) It wasn't a good look for a company that prides itself on working closely with the academic community\. ![Anthropic's Claude mythos log is shown](https://www.engadget.com/img/gallery/anthropic-backtracks-on-policy-that-sabotaged-researchers-work/intro-1781168344.jpg)Primakov/Shutterstock Anthropic is walking back a policy that discreetly hamstrung researchers using its new[Claude Fable 5 LLM](https://www.engadget.com/2190934/anthropic-fable-ai-brings-the-capabilities-of-its-unreleased-mythos-model-to-regular-users/)to create competing AI models, the company told*[Wired](https://www.wired.com/story/anthropic-responds-to-backlash-on-claudes-secret-sabotage-on-ai-research/)*\. "We're changing Fable 5's safeguards for frontier LLM development to make them visible," the company said in a statement\. "We made the wrong tradeoff and we apologize for not getting the balance right\." When Anthropic released Claude Fable 5, a new model based on its powerful Mythos system, researchers noted something odd\. They found that that Fable 5 would quietly reroute requests to a lesser model when asked to perform certain actions\. Moreover, that restriction wasn't disclosed in the model's documentation\. The new model was either refusing or degrading responses for tasks like training competing LLMs, debugging AI code and optimizing neural architecture\. Researchers were bothered not only by that degradation but by Anthropic's lack of transparency about it\. They were also concerned, of course, that they had burned tokens and money for a model that didn't do what they expected\. Anthropic has painted itself as a more ethical and researcher\-friendly alternative to OpenAI, so its actions with Fable 5 created a swift backlash\. "Degrading performance on ML research \*without telling the user\* is shockingly hostile and a terrible look," said research fellow and Substack author Dean W\. Ball[on X](https://x.com/deanwball/status/2064434861088395730)\. Anthropic isn't reversing its safeguard policy on Fable 5, but rather making the restrictions visible to users\. "If the company suspects a user is trying to use Claude to build a highly capable AI it will alert them that it's either refusing the request, or rerouting the user to a less capable model,"*Wired*wrote\.

Similar Articles