risk-mitigation

Tag

Cards List
#risk-mitigation

How are you actually building approval gates for agents? I'm convinced most are meaningless rubber stamps

Reddit r/AI_Agents · 2026-06-23

The author argues that many human approval gates for AI agents are ineffective rubber stamps, and proposes a framework for designing meaningful review mechanisms that actually catch errors.

0 favorites 0 likes
#risk-mitigation

@levie: The layer that can route to the best AI model for the particular job is going to increase in value substantially. There…

X AI KOLs Following · 2026-06-14 Cached

A tweet argues that the layer routing between AI models will become increasingly valuable due to cost optimization, capability differences, and risk mitigation, while quoting OpenRouter's Fusion API announcement.

0 favorites 0 likes
#risk-mitigation

I built a plugin that makes OpenClaw ask my phone before doing anything risky

Reddit r/openclaw · 2026-06-11

OKed is a plugin for OpenClaw that intercepts risky tool calls and requires user approval before execution, preventing agents from performing destructive actions like deleting data or sending payments.

0 favorites 0 likes
#risk-mitigation

Our commitment to community safety

OpenAI Blog · 2026-04-28 Cached

OpenAI outlines its commitment to community safety, detailing how ChatGPT is trained to detect and mitigate risks of violence and harm through refined safeguards and expert input.

0 favorites 0 likes
#risk-mitigation

Preparing for future AI risks in biology

OpenAI Blog · 2025-06-18 Cached

OpenAI publishes a comprehensive approach to managing dual-use risks from advanced AI models in biology, outlining strategies for enabling beneficial scientific discovery while preventing misuse for bioweapons development through expert collaboration, model training, detection systems, and security controls.

0 favorites 0 likes
#risk-mitigation

Updating the Frontier Safety Framework

Google DeepMind Blog · 2025-02-04 Cached

DeepMind has published an updated Frontier Safety Framework (v2.0) with stronger security protocols for frontier AI models, including new Critical Capability Level (CCL) security recommendations and enhanced approaches to deceptive alignment risks. The framework aims to prevent unauthorized model weight exfiltration and manage risks as AI systems become more powerful.

0 favorites 0 likes
#risk-mitigation

OpenAI’s Approach to Frontier Risk

OpenAI Blog · 2023-10-26 Cached

OpenAI publishes details on its approach to frontier AI risks and announces progress on voluntary safety commitments made in July 2023, including the release of DALL-E 3 system card and the development of a new Preparedness Framework to manage catastrophic risks from advanced AI systems.

0 favorites 0 likes
← Back to home

Submit Feedback