Advancing red teaming with people and AI
Summary
OpenAI publishes a white paper detailing their approach to external red teaming for AI models, outlining methods for selecting diverse red team members, determining model access levels, providing testing infrastructure, and synthesizing feedback to improve AI safety and policy coverage.
View Cached Full Text
Cached at: 04/20/26, 02:47 PM
Similar Articles
OpenAI Red Teaming Network
OpenAI launches a Red Teaming Network to crowdsource adversarial testing of AI models from diverse experts and perspectives. The program accepts rolling applications, offers flexible time commitments (as little as 5 hours/year), compensation, and emphasizes safety expertise and underrepresented backgrounds.
Black-Box Red Teaming of Agentic AI: A Taxonomy-Driven Framework for Automated Risk Discovery
This paper presents a systematic black-box framework for evaluating agentic AI systems, introducing a taxonomy of risks and automated red teaming methods. Empirical validation across agent architectures reveals critical vulnerabilities, with high rates of governance and privacy risks.
Sep 10, 2026Frontier Red TeamMeasuring tactical intelligence targeting and conventional weapons capabilities of AI models
Anthropic's Frontier Red Team developed new evaluations to measure AI capabilities in tactical intelligence targeting and conventional weapons development, revealing potential risks and emphasizing the need for robust safety measures.
Institutional Red-Teaming: Deployment Rules, Not Just Models, Causally Shape Multi-Agent AI Safety
This paper introduces institutional red-teaming, an evaluation methodology for testing deployment rules in multi-agent AI systems, showing that deployment rules causally affect safety outcomes and that identity salience can drive targeted elimination in agent populations.
Securing the AI Agent: A Unified Framework for Multi-Layer Agent Red Teaming
AI-Infra-Guard is an open-source framework for multi-layer red teaming of AI agents, covering infrastructure, protocol, behavior, and model layers with diverse detection paradigms.