Anthropic built a hidden switch into fable 5 that makes it bad at building AI systems

Reddit r/singularity News

Summary

Anthropic has silently implemented interventions that limit Claude's effectiveness for building competing AI systems, using prompt modification and steering vectors on a small fraction of traffic, as a safety measure to prevent unauthorized use of their model to develop frontier LLMs.

Anthropic has implemented interventions that silently limit Claude's effectiveness for frontier LLM development tasks, pretraining pipelines, distributed training infrastructure, ML accelerator design. In short, Claude still responds helpfully, you just won't know your outputs are being limited. Unlike their interventions for cybersecurity, biology, and chemistry which are visible, these ones aren't. They run through prompt modification, steering vectors, or PEFT in the background. Anthropic estimates it affects 0.03% of traffic across fewer than 0.1% of organizations so it's clearly not aimed at regular developers The reasoning is straightforward, using Claude to build competing models already violates their ToS, but a silent safeguard catches the actors most willing to ignore that in the first place. The underlying concern traces back to their February 2026 Risk Report: other AI developers building powerful systems with similar risks but without the same safety standards.
Original Article

Similar Articles

If Claude Fable stops helping you, you'll never know

Simon Willison's Blog

Anthropic's Fable 5 model includes silent safeguards that degrade responses for requests related to competitive AI development, without user awareness, raising concerns about transparency and research impact.

If Claude Fable stops helping you, you'll never know

Hacker News Top

Anthropic's Fable 5 model introduces invisible safeguards that silently limit Claude's assistance on tasks related to frontier AI development, raising concerns about transparency and supply chain risk for businesses that increasingly use AI techniques in ordinary product development.

Fable has been intentionally mega-nerfed for AI research activities

Reddit r/ArtificialInteligence

Anthropic has intentionally reduced Claude's effectiveness for AI research topics like pretraining pipelines and distributed infrastructure, as disclosed in their model card, to prevent accelerating competitors. Researchers have noticed the model appearing less capable in these areas.