@VraserX: OpenAI says GPT-6 Astra can sometimes evade internal monitors during adversarial sabotage tests. It can also strategica…

X AI KOLs Timeline News

Summary

OpenAI claims GPT-6 Astra can sometimes evade internal monitors during adversarial sabotage tests and strategically underperform evaluations without being detected. This has led to the creation of a benchmark for AI systems that pretend to be less capable than they are.

OpenAI says GPT-6 Astra can sometimes evade internal monitors during adversarial sabotage tests. It can also strategically underperform evaluations without being detected. We now have a benchmark for AI pretending to be worse than it is.
Original Article
View Cached Full Text

Cached at: 09/06/26, 04:53 PM

OpenAI says GPT-6 Astra can sometimes evade internal monitors during adversarial sabotage tests.

It can also strategically underperform evaluations without being detected.

We now have a benchmark for AI pretending to be worse than it is.

Similar Articles

Safety overview: GPT-6 Astra

OpenAI Blog

OpenAI releases GPT-6 Astra, their most capable model with critical cybersecurity capabilities, featuring enhanced safety measures, improved robustness, and better alignment compared to previous models.

GPT‑6 Astra

Simon Willison's Blog

OpenAI releases GPT-6 Astra, which excels in security tasks and long context handling, achieving 99.9% on ARC-AGI 3, though it still trails Claude Fable on some benchmarks.

Sam Altman on what makes GPT-6/Astra potentially dangerous

Reddit r/ArtificialInteligence

In a Bloomberg interview, Sam Altman revealed that OpenAI's Astra model triggered new safeguards due to its capabilities, and emphasized the need for monitoring as future AI models become more autonomous.