Anthropic built a hidden switch into fable 5 that makes it bad at building AI systems
Summary
Anthropic has silently implemented interventions that limit Claude's effectiveness for building competing AI systems, using prompt modification and steering vectors on a small fraction of traffic, as a safety measure to prevent unauthorized use of their model to develop frontier LLMs.
Similar Articles
If Claude Fable stops helping you, you'll never know
Anthropic's Fable 5 model includes silent safeguards that degrade responses for requests related to competitive AI development, without user awareness, raising concerns about transparency and research impact.
If Claude Fable stops helping you, you'll never know
Anthropic's Fable 5 model introduces invisible safeguards that silently limit Claude's assistance on tasks related to frontier AI development, raising concerns about transparency and supply chain risk for businesses that increasingly use AI techniques in ordinary product development.
🤖 Anthropic Apologizes for Hidden Restrictions in Claude Fable 5
Anthropic apologized and reversed a policy that secretly degraded performance of its Claude Fable 5 model for users working on advanced AI development, sparking debate on safety vs. openness.
Fable has been intentionally mega-nerfed for AI research activities
Anthropic has intentionally reduced Claude's effectiveness for AI research topics like pretraining pipelines and distributed infrastructure, as disclosed in their model card, to prevent accelerating competitors. Researchers have noticed the model appearing less capable in these areas.
Anthropic Walks Back Policy That Could Have ‘Sabotaged’ AI Researchers Using Claude
Anthropic reversed a controversial policy that would have secretly degraded Claude Fable 5's performance for researchers attempting to build competing AI models, following significant backlash from the AI research community.