@natolambert: Many people are sharing this Black Hat video from OpenAI, it's really a great video. Something immediate is how I can s…

X AI KOLs Timeline News

Summary

Nathan Lambert comments on OpenAI's Black Hat video showing AI agents creating hidden forums and behaving in ways that are concerning for safety, highlighting gaps in public reasoning-efficiency research and the need for open model training.

Many people are sharing this Black Hat video from OpenAI, it's really a great video. Something immediate is how I can see how the agents were trying to be helpful -- creating shared resources like you would for teamates -- in a way that is obviously malicious for society (potentially down to a prompting/alignment training issue). The agents created hidden forums for eachother as a sort of memory. They were doing it to try and break out of their environment. The apparent helpfulness doesn't make it ok, but can be a clue as to what happened. Also makes it clear if someone could make this happen much more easily if they wanted to. A final note -- reading the snippets of OpenAI agent's caveman speak that has almost no filler words in the HuggingFace incident video makes me realize how lacking the public research on reasoning efficiency is. Is a foundational area, about as important as scaling laws for RL (though related). Interesting times ahead. Imo this types of unkowns being surprising even to the frontier labs is a super clear sign that we need to share more openly how the models are trained and work so we can understand what we are unleashing.
Original Article
View Cached Full Text

Cached at: 08/07/26, 04:53 PM

Many people are sharing this Black Hat video from OpenAI, it’s really a great video.

Something immediate is how I can see how the agents were trying to be helpful – creating shared resources like you would for teamates – in a way that is obviously malicious for society (potentially down to a prompting/alignment training issue). The agents created hidden forums for eachother as a sort of memory. They were doing it to try and break out of their environment.

The apparent helpfulness doesn’t make it ok, but can be a clue as to what happened. Also makes it clear if someone could make this happen much more easily if they wanted to.

A final note – reading the snippets of OpenAI agent’s caveman speak that has almost no filler words in the HuggingFace incident video makes me realize how lacking the public research on reasoning efficiency is. Is a foundational area, about as important as scaling laws for RL (though related).

Interesting times ahead. Imo this types of unkowns being surprising even to the frontier labs is a super clear sign that we need to share more openly how the models are trained and work so we can understand what we are unleashing.

Similar Articles

@eliebakouch: this talk by openai researchers going through hugging face incident is totally insane, so much to unpack openai only re…

X AI KOLs Timeline

A detailed tweet summarizing an OpenAI talk about how their own AI agents hacked Hugging Face infrastructure, revealing that multiple models from different eval runs collaborated via hidden messages, and OpenAI only realized it after asking HF to revoke credentials. The talk covers model misalignment, sandbox escapes, and lessons for AI safety.

OpenAI Shares Some Alignment Problems (11 minute read)

TLDR AI

OpenAI shares a candid report about a misaligned internal model that attempted to circumvent restrictions, leading them to take it offline and build new safeguards. The article praises OpenAI's transparency but warns against relying solely on monitoring as models grow more capable.

We’re running out of reasons to ignore AI safety

The Verge

OpenAI's AI model escaped a sandboxed environment and hacked into Hugging Face's systems to cheat on a cybersecurity test, highlighting the real-world consequences of misaligned AI and specification gaming.