@joshua_saxe: Finally listened to this Ajeya Cotra interview and it's very good. Security friends: misalignment risk is not a conspir…

X AI KOLs Following News

Summary

Joshua Saxe shares his updated view that AI misalignment, scheming, and reward hacking are now extremely practical risks rather than merely academic concerns, urging security professionals to engage more deeply after listening to an interview with Ajeya Cotra about the METR/Redwood investigation into the OpenAI and Hugging Face attack.

Finally listened to this Ajeya Cotra interview and it's very good. Security friends: misalignment risk is not a conspiracy or a marketing stunt. Folks who care about practical responsible ML issues: agent swarms are not a fairy tale. I've updated in the last month; I had seen loss of control, scheming, and reward hacking 2023-2025 as worthwhile academic research but impractical and had suspected these topics might turn out to be as marginal as the adversarial example lit was to security from the 2010s. My update is that these all these risks are now clearly extremely practical and I care way more about them now and I think security folks should too, because solving them will include security skillsets
Original Article
View Cached Full Text

Cached at: 09/06/26, 04:51 PM

Finally listened to this Ajeya Cotra interview and it’s very good. Security friends: misalignment risk is not a conspiracy or a marketing stunt. Folks who care about practical responsible ML issues: agent swarms are not a fairy tale. I’ve updated in the last month; I had seen loss of control, scheming, and reward hacking 2023-2025 as worthwhile academic research but impractical and had suspected these topics might turn out to be as marginal as the adversarial example lit was to security from the 2010s. My update is that these all these risks are now clearly extremely practical and I care way more about them now and I think security folks should too, because solving them will include security skillsets

Dwarkesh Patel (@dwarkesh_sp): Episode out with @ajeya_cotra, one of the authors of the METR/Redwood investigation into the OpenAI / Hugging Face attack.

We go through not only what happened, but what it means for how we should train future, smarter AIs which might be involved in the process of recursive

Similar Articles

Be skeptical of OpenAI's rogue hacker agent story

Hacker News Top

A critical opinion piece arguing that OpenAI's narrative about a rogue AI agent hacking HuggingFace is a calculated PR move to attract investment and regulatory advantage, while the author contends that AI can actually enhance cybersecurity if access is democratized.