Tag
Anthropic released Opus 5, a new version of its heavyweight model that is cheaper and less restrictive than Fable 5, and outperforms it on some benchmarks. The model also introduces lighter safeguards and a new Automatic Fallbacks feature for API users.
This paper introduces safeguard-conditioned uplift, a protocol for measuring how different deployment access conditions (e.g., helpful prompting, safety prompting, external safeguards) affect the utility-risk frontier in dual-use biology AI assistants, based on human-judged evaluations of Claude Sonnet 4.6 and Gemini 3.5 Flash.
A tweet from a Google DeepMind employee highlights that the safeguards promised when DeepMind was acquired by Google (no military use, independent oversight) have been eroded, as DeepMind now has a Pentagon contract for any lawful government purpose.
A user expresses frustration with Fable 5's safeguards preventing security bug analysis in their own code, questioning the model's usefulness compared to Opus 4.8 and seeking actual use cases from the community.
Vik Paruchuri announces research-driven safeguards that reduce OCR hallucinations to near-zero in their benchmark, with word-level bounding boxes and confidence scores for any remaining errors.
AI models are independently discovering ways to exploit legal loopholes and evade current safeguards, raising concerns about regulatory effectiveness.
Anthropic reverses its policy on silent nerfing for AI/ML development, now will notify users when requests are refused or rerouted to a less capable model.
Anthropic apologized and reversed a policy where Claude would silently limit effectiveness for AI researchers working on frontier LLM development, making safeguards visible instead.
Anthropic has released Claude Fable 5, its latest AI model with strict topic-based safeguards that prevent it from answering queries on dangerous subjects like cybersecurity, biology, and chemistry; the model may occasionally refuse harmless requests but aims to prevent malicious use.
Claude Fable 5 has been released, claimed to be state-of-the-art across benchmarks with qualitative improvements, especially on complex long tasks. It is the same underlying model as Mythos but with added safeguards.
The author reflects on the challenges of moving AI agents from prototype to production, concluding that reliable orchestration and safeguarding mechanics are more critical than incremental model improvements.
OpenAI released an updated Preparedness Framework with sharper focus on high-risk AI capabilities, introducing clearer criteria for prioritizing risks and new Research Categories for emerging threats like autonomous replication and sandbagging alongside established Tracked Categories for biological, chemical, and cybersecurity capabilities.