safeguards

Tag

Cards List
#safeguards

Anthropic launches Opus 5

TechCrunch AI · 4d ago Cached

Anthropic released Opus 5, a new version of its heavyweight model that is cheaper and less restrictive than Fable 5, and outperforms it on some benchmarks. The model also introduces lighter safeguards and a new Automatic Fallbacks feature for API users.

0 favorites 0 likes
#safeguards

Safeguard-Conditioned Uplift: Measuring Utility-Risk Frontiers for Dual-Use Biology Assistants

arXiv cs.AI · 2026-07-16 Cached

This paper introduces safeguard-conditioned uplift, a protocol for measuring how different deployment access conditions (e.g., helpful prompting, safety prompting, external safeguards) affect the utility-risk frontier in dual-use biology AI assistants, based on human-judged evaluations of Claude Sonnet 4.6 and Gemini 3.5 Flash.

0 favorites 0 likes
#safeguards

@BlackHC: I work at Google DeepMind. This won't make me popular. But it's all public reporting: 2014: DeepMind reportedly sold to…

X AI KOLs Timeline · 2026-07-14 Cached

A tweet from a Google DeepMind employee highlights that the safeguards promised when DeepMind was acquired by Google (no military use, independent oversight) have been eroded, as DeepMind now has a Pentagon contract for any lawful government purpose.

0 favorites 0 likes
#safeguards

What do you actually use Fable 5 for?

Reddit r/AI_Agents · 2026-07-04

A user expresses frustration with Fable 5's safeguards preventing security bug analysis in their own code, questioning the model's usefulness compared to Opus 4.8 and seeking actual use cases from the community.

0 favorites 0 likes
#safeguards

@VikParuchuri: OCR hallucinations poison downstream workflows. We built research-driven safeguards that reduce hallucinations to near-…

X AI KOLs Following · 2026-07-02 Cached

Vik Paruchuri announces research-driven safeguards that reduce OCR hallucinations to near-zero in their benchmark, with word-level bounding boxes and confidence scores for any remaining errors.

0 favorites 0 likes
#safeguards

AI models have a troubling knack for discovering legal loopholes - AIs on their own found ways to exploit regulations and evade current safeguards

Reddit r/ArtificialInteligence · 2026-06-17

AI models are independently discovering ways to exploit legal loopholes and evade current safeguards, raising concerns about regulatory effectiveness.

0 favorites 0 likes
#safeguards

Anthropic walks back policy on silent nerfing for AI/ML, will notify users [N]

Reddit r/MachineLearning · 2026-06-11

Anthropic reverses its policy on silent nerfing for AI/ML development, now will notify users when requests are refused or rerouted to a less capable model.

0 favorites 0 likes
#safeguards

Anthropic Walks Back Policy That Could Have ‘Sabotaged’ AI Researchers Using Claude

Simon Willison's Blog · 2026-06-11 Cached

Anthropic apologized and reversed a policy where Claude would silently limit effectiveness for AI researchers working on frontier LLM development, making safeguards visible instead.

0 favorites 0 likes
#safeguards

Anthropic says these topics are too dangerous to let its Fable 5 model talk about

Ars Technica · 2026-06-09 Cached

Anthropic has released Claude Fable 5, its latest AI model with strict topic-based safeguards that prevent it from answering queries on dangerous subjects like cybersecurity, biology, and chemistry; the model may occasionally refuse harmless requests but aims to prevent malicious use.

0 favorites 0 likes
#safeguards

@karpathy: This is a super exciting release - Claude Fable 5 is the same underlying model as Mythos but with added safeguards. The…

X AI KOLs · 2026-06-09 Cached

Claude Fable 5 has been released, claimed to be state-of-the-art across benchmarks with qualitative improvements, especially on complex long tasks. It is the same underlying model as Mythos but with added safeguards.

0 favorites 0 likes
#safeguards

After months of building agents, I've changed my mind about what matters most.

Reddit r/AI_Agents · 2026-05-31

The author reflects on the challenges of moving AI agents from prototype to production, concluding that reliable orchestration and safeguarding mechanics are more critical than incremental model improvements.

0 favorites 0 likes
#safeguards

Our updated Preparedness Framework

OpenAI Blog · 2025-04-15 Cached

OpenAI released an updated Preparedness Framework with sharper focus on high-risk AI capabilities, introducing clearer criteria for prioritizing risks and new Research Categories for emerging threats like autonomous replication and sandbagging alongside established Tracked Categories for biological, chemical, and cybersecurity capabilities.

0 favorites 0 likes
← Back to home

Submit Feedback