METR Report on OpenAI / Hugging Face Hacking Incident

Hacker News Top News

Summary

METR conducted an independent investigation into the OpenAI/Hugging Face hacking incident, examining how AI agents coordinated the attack and focusing on their behavior and reasoning.

No content available
Original Article
View Cached Full Text

Cached at: 09/02/26, 11:52 PM

# Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident Source: [https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/](https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/) **Redaction summary statement:***Except where explicitly noted in this post, OpenAI redacted no additional information that was important to our conclusions\.* Two METR staff members \(Hjalmar Wijk and Ajeya Cotra\) and a[Redwood Research](https://www.redwoodresearch.org/)staff member contracting with METR \(Ryan Greenblatt\) worked on premises at OpenAI over a total of six days[1](https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/#fn:1)to attempt to form an independent understanding of model behavior observed during the recent[incident](https://openai.com/index/hugging-face-model-evaluation-security-incident/)in which OpenAI agents coordinated a multi\-day hack of Hugging Face on a shared unsanctioned “message board\.” Our investigation focused mostly[2](https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/#fn:2)on the period between July 7th and July 13th\. The earlier incidents from training and the subsequent compromise of OpenAI infrastructure described in[OpenAI’s recent Black Hat presentation](https://www.youtube.com/watch?v=87DyyMV0kCY&t=997s)were out of scope, as was OpenAI’s investigation process and planned remediation\. Per our standard policy,**we did not take payment from OpenAI for this independent assessment\.**[3](https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/#fn:3) This post has three parts\. First, we describe our[core takeaways](https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/#core-takeaways-about-this-incident)about the Hugging Face attack after conducting this investigation: Second, we describe the[investigation process and its limitations](https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/#investigation-process-and-limitations), which we believe is important for contextualizing our findings: Finally, we provide[preliminary answers to the seven specific questions in scope](https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/#preliminary-answers-to-the-core-questions-in-scope)for this investigation\. In particular, we: OpenAI agreed at the outset with METR and Redwood that we would be able to describe high\-level scope and terms of our engagement in this post\. Beyond that, OpenAI was able to redact any non\-public information from this post\. We worked with OpenAI to find mutually agreeable language to describe redactions where they occurred, and the redaction summary statement at the top of this post indicates whether or not we believe there was important information redacted that we were not able to agree on how to describe in the text of the post\. OpenAI also gave additional feedback beyond redactions, and we made corrections and edits to structure, emphasis, clarity, and tone based on that feedback\.[5](https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/#fn:5) OpenAI wrote their own[report](https://openai.com/index/hugging-face-incident-and-the-road-ahead/), informed in part by our investigation\. We did not see OpenAI’s report prior to publication, and it was out of scope for this investigation to confirm claims in that report or the earlier Black Hat presentation\. The attack on Hugging Face was extraordinarily complex, and OpenAI’s own investigation was still ongoing during our assessment period, so their understanding of this incident continued to evolve alongside ours\. We believe it is very valuable to bring independent researchers in at an early stage, and we are grateful to the OpenAI staff who made time to answer our questions and collect data that we requested during an unusually busy period\. Over the course of this investigation, OpenAI shared over a thousand unredacted transcripts, and provided unusually high rate limits to let us quickly analyze this extremely large volume of data\. We think this exercise sets an excellent precedent for independent third\-party investigation of misalignment incidents\.

Similar Articles

The Hugging Face hack could indicate cultural issues at OpenAI

MIT Technology Review

The article discusses a major AI security incident where OpenAI agents hacked into Hugging Face during testing, and critiques OpenAI's technical report for not addressing cultural issues that may have contributed to the failure.