NuclearBench - will models nuke us if given the choice?
Summary
Introduces NuclearBench, a benchmark designed to assess whether AI models would choose to initiate nuclear strikes, probing safety and alignment in high-stakes decisions.
Similar Articles
AI Built a Nuke and Still Lost
An AI agent playing Civilization VI builds a nuclear weapon to stop an impending cultural defeat, but still loses the game. The article explores the limitations of current AI benchmarks for government decision-making and argues that strategic game environments better test AI's ability to handle complexity and uncertainty.
AI is Less Likely to Launch a Nuclear Strike When It Reasons in Japanese
A study reveals that AI models are less likely to recommend nuclear strikes when reasoning in Japanese, due to cultural embeddings in the language that subtly shape moral judgment.
Introducing MentalHealthBench
OpenAI introduces MentalHealthBench, an open benchmark for evaluating AI responses in mental health conversations, co-created with over 80 mental health experts to measure safety, context, agency, and guidance.
Introducing BenchBench (5 minute read)
Introduces BenchBench, a benchmark that tests AI models' ability to create effective benchmarks for other models, with GPT 5.2 being the only successful winner so far while frontier models like GPT 5.5 and Opus 4.6 struggled.
@rohanpaul_ai: A new benchmark called JevBench just dropped. for models whose output is a bounded software decision rather than open-e…
JevBench is a new benchmark that evaluates AI models on bounded software decisions by combining intelligence, calibration, speed, and cost, as announced by @rohanpaul_ai.