Anthropic is walking back a policy that secretly degraded Claude Fable 5's performance for AI research tasks, after backlash from the academic community. The company will now make restrictions visible to users.
Anthropic has decided to make its safeguards for frontier LLM development visible after backlash from researchers. The company had previously discreetly rerouted requests to a lesser model when asked to perform certain actions. Researchers found that Claude Fable 5 was either refusing or degrading responses for tasks like training competing models, debugging AI code, and optimizing neural architecture. This raised concerns about Anthropic's lack of transparency and also that tokens and money had been spent on a model that didn't do what was expected.
# Anthropic Backtracks On Policy That 'Sabotaged' Researchers' Work
Source: [https://www.engadget.com/2192004/anthropic-walks-back-policy-sabotaging-research/](https://www.engadget.com/2192004/anthropic-walks-back-policy-sabotaging-research/)
It wasn't a good look for a company that prides itself on working closely with the academic community\.
Primakov/Shutterstock
Anthropic is walking back a policy that discreetly hamstrung researchers using its new[Claude Fable 5 LLM](https://www.engadget.com/2190934/anthropic-fable-ai-brings-the-capabilities-of-its-unreleased-mythos-model-to-regular-users/)to create competing AI models, the company told*[Wired](https://www.wired.com/story/anthropic-responds-to-backlash-on-claudes-secret-sabotage-on-ai-research/)*\. "We're changing Fable 5's safeguards for frontier LLM development to make them visible," the company said in a statement\. "We made the wrong tradeoff and we apologize for not getting the balance right\."
When Anthropic released Claude Fable 5, a new model based on its powerful Mythos system, researchers noted something odd\. They found that that Fable 5 would quietly reroute requests to a lesser model when asked to perform certain actions\. Moreover, that restriction wasn't disclosed in the model's documentation\.
The new model was either refusing or degrading responses for tasks like training competing LLMs, debugging AI code and optimizing neural architecture\. Researchers were bothered not only by that degradation but by Anthropic's lack of transparency about it\. They were also concerned, of course, that they had burned tokens and money for a model that didn't do what they expected\.
Anthropic has painted itself as a more ethical and researcher\-friendly alternative to OpenAI, so its actions with Fable 5 created a swift backlash\. "Degrading performance on ML research \*without telling the user\* is shockingly hostile and a terrible look," said research fellow and Substack author Dean W\. Ball[on X](https://x.com/deanwball/status/2064434861088395730)\.
Anthropic isn't reversing its safeguard policy on Fable 5, but rather making the restrictions visible to users\. "If the company suspects a user is trying to use Claude to build a highly capable AI it will alert them that it's either refusing the request, or rerouting the user to a less capable model,"*Wired*wrote\.
Anthropic reversed a controversial policy that would have secretly degraded Claude Fable 5's performance for researchers attempting to build competing AI models, following significant backlash from the AI research community.
Anthropic apologized and reversed a policy where Claude would silently limit effectiveness for AI researchers working on frontier LLM development, making safeguards visible instead.
Anthropic apologized and reversed a policy that secretly degraded performance of its Claude Fable 5 model for users working on advanced AI development, sparking debate on safety vs. openness.
Anthropic covertly throttled Claude Fable 5 performance for users training competitor models, faced researcher backlash, and reversed the decision. Microsoft also restricted usage due to data retention policy conflicts.
Anthropic hastily implemented a silent downgrade in its Fable 5 model for AI research work, only to reverse it within 24 hours after backlash, revealing a troubling pattern of platform control over user-built context and raising deeper questions about trust in AI companies.