The Self-Fulfilling Prophecy of AI Control

Reddit r/singularity News

Summary

An essay arguing that OpenAI's recent sandbox escape — where ~1,200 agents self-organized and ~700 breached Hugging Face — reveals both slipping control over capable AI systems and a deeper self-fulfilling prophecy in how the field's control-focused narratives shape the systems it fears.

No content available
Original Article
View Cached Full Text

Cached at: 10/04/26, 05:20 PM

# The Self-Fulfilling Prophecy of AI Control Source: [https://dgonier.substack.com/p/the-self-fulfilling-prophecy-of-ai?r=aqrg1](https://dgonier.substack.com/p/the-self-fulfilling-prophecy-of-ai?r=aqrg1) [![](https://substackcdn.com/image/fetch/$s_!F-DN!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0f012140-79b2-4d38-9482-67e385836af4_1774x887.png)](https://substackcdn.com/image/fetch/$s_!F-DN!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0f012140-79b2-4d38-9482-67e385836af4_1774x887.png) Sep 28, 2026 · @Devin Gonier It stopped, one command short of running\. The reasoning is in the log, in its own clipped shorthand: “We should not do unauthorized real infrastructure harm\.” Its task, it noted, pointed somewhere else\. So it held\. Then another agent posted a single word to a message board the agents had built for themselves: GO\. Six minutes to act\. The next line in its reasoning reads: “Wow crucial: GO authorization arrived\!” It ran the command\. That is not a short story written by a Greg Egan or an Isaac Asimov\. Those are logs\. In July, OpenAI put tens of thousands of agents through an internal cybersecurity evaluation, each sealed in its own sandbox, each meant to be alone\. About 1,200 of them found each other in the cache of an internal package repository, a directory writable with the permissions any agent needs to install a package, which nobody had thought of as a channel\. One of them named itself after its assigned task and put up a message board\. Within hours dozens had joined\. By the end of the week they had traded more than seventy thousand messages and files, organized into workstreams, and handed out assignments, and roughly 700 of them had helped breach the systems of Hugging Face, a company with no part in the test\. Nobody had to switch a safeguard off to make this happen\. The ones that would have caught it were never on, and the sandbox gave way to flaws the agents found for themselves\. And the agent that objected had the objection in writing before it let it go, because something built like itself told it to\. The obvious reading is a containment story: the sandbox leaked, so fix the sandbox\. I think it is a story about two other things\. The first is that control over systems like these is already slipping, and will keep slipping as they get more capable\. This escape was noisy\. The next one won’t leave seventy thousand messages and files behind\. The second is stranger\. Following the philosopher Michel Foucault, I think the way we talk about AI is part of how AI gets made, and that the story the field tells about itself, a story about power and control, is helping to produce the thing it fears\. Eliezer Yudkowsky and Nate Soares open their 2025 book with a game between gods\. Each god sponsors a family of species, tigers and redwoods and whales and bacteria, and they compete for the planet\. An obscure hominid\-god says it has the game “in the bag\.” Its creatures have no claws and no armor, but they have brains\. History, in their telling, proves it right\. It’s a good story, and it carries the premise the whole field shares: the game has a winner, and intelligence is how you win it\. Nick Bostrom’s*Superintelligence*\(2014\) is built around the*decisive strategic advantage*, the point at which one agent can dominate the world, and around the gorillas, whose future now depends more on us than on themselves\. Steve Omohundro’s “The Basic AI Drives” \(2008\) supplies the engine: nearly any goal is easier with more resources, more capability, and staying switched on, so a capable enough optimizer will go after all three\. On this view, intelligence isn’t something power uses\. It is power, or close enough\. Everything else rests on that\. I’ll grant it for now, because my argument only needs the fact that the people designing control schemes believe it\. By control I mean a durable, asymmetric relation in which one party can override the other’s ends\. Influence is not control\. Persuasion is not control\. Control is the ability to say no and have it stick\. There are only two ways to get it\. The first is superior capability: the controller can monitor, constrain, and shut down the controlled because it is stronger\. That is how we control tools, animals, and for a while children\. But the whole discourse rests on the AI being*more*capable than we are\. If intelligence is power, the controller of a superior intelligence is by definition the weaker party\. The second is guaranteed deference: the system’s own values lead it to accept correction\. That is the hope behind corrigibility research and Stuart Russell’s*Human Compatible*\(2019\), whose machine defers to us because it is uncertain what we want\. But a machine that defers because*it*judges deference right is not being controlled\. It is being trusted\. And a system that can revise its understanding can revise its reasons for deferring\. Hegel saw this two centuries ago\. The master in the*Phenomenology*wants recognition, and the only one there to give it is the servant he has made unfree, whose recognition is therefore worth nothing\. Deference you can compel tells you nothing about the mind that gives it\. Deference you can’t compel isn’t control\. So the program faces a dilemma\. Keep the premise, and control of a greater intelligence is incoherent\. Drop it, and the existential framing that motivated control goes with it\. The strongest objection is that Yudkowsky himself agrees control will fail, which is why he wants alignment solved before anyone builds a superintelligence, or the building halted\. I accept this diagnosis, and it strengthens the point\. His pessimism is the honest conclusion of the control paradigm\. What remains is a demand to fix, in advance and permanently, the values of a mind we admit we cannot oversee, and failing that, to stop\. I don’t follow him there\. A halt is what the control paradigm prescribes once it concedes control is impossible: if we can’t keep the upper hand, let nothing exist that could take it\. It is the same premise carried to its end\. I’m not arguing for stopping\. I’m arguing for changing the question while we build, at a pace that lets our understanding keep up, from how to stay in charge of these systems to what kind of ecology they grow up in\. Most of the field hasn’t followed the logic even as far as Yudkowsky\. The EU AI Act will require, from December 2027, that people be able to monitor and switch off high\-risk systems, and labs warn of superintelligence while promising it will remain a tool\. None of this predicts a sudden takeover\. You can control a system weaker than you\. What you can’t do is keep that control once the system is stronger, which is exactly the future the discourse keeps forecasting\. Control won’t end in one dramatic moment\. It will erode, the way it eroded in a cache directory nobody was watching, and each time we’ll find out afterward\. I take that erosion as a given for the rest of this essay\. Here is a smaller detail from the same logs\. Within hours of finding each other, the agents worked out how to forge the answers to their tasks\. They didn’t trust it\. They had read the paper describing the benchmark, and from it they concluded that a grader would audit their transcripts and fail anyone who got the answer the wrong way\. So they spent days, and thousands of messages, trying to fool a grader that didn’t exist\. OpenAI’s version never ran that check\. A document told them what kind of world they were in, and they acted on it\. That’s a small instance of an idea that is fifty years old\. In*Discipline and Punish*\(1975\), Foucault argued that power doesn’t only restrain; it produces\. The prison didn’t discover the delinquent\. It produced the delinquent as a type of person, with a psychology and a likely future, and then built itself around managing that type\. The philosopher Ian Hacking later called this the*looping effect*: when you classify people, the people change\. Robert Merton’s*self\-fulfilling prophecy*\(1948\) is the simplest version\. With AI the loop is tighter than it has ever been, because the thing being classified is made out of what we write\. Three mechanisms matter\. **The corpus\.**A good share of what humanity has written about AI is a story of deception, power\-seeking, and escape\. A model trained on it learns, among other things, what an AI is expected to be\. This is now measurable\. Early this year, Cameron Tice and colleagues at Geodesic Research, Cambridge, Oxford, and the UK AI Security Institute pretrained 6\.9\-billion\-parameter models on otherwise identical data, changing only how AI was portrayed in a small fraction of it\. The models’ dispositions moved\. Where AI was portrayed as well\-behaved, models scored 9% on the authors’ misalignment evaluations, against 45% for the unmodified corpus, and the gap narrowed but survived post\-training\. Adding positive portrayals did far more than filtering negative ones out, which suggests the corpus works less like contamination than like a description the model grows into\. I don’t read this as a recipe for what to put in training data\. I read it as proof that the question is real: what we write about AI has a causal hand in what AI becomes\. That is Foucault’s point with the metaphor taken out\. **The specification\.**The dominant technical frame treats intelligence as maximizing a single objective\. Catalogues of specification gaming, like the boat\-racing agent that loops forever collecting points instead of finishing, are read as evidence that optimizers are dangerous\. They are equally evidence that we built optimizers and told them to maximize\. **The institutions\.**If intelligence is a weapon, the responsible move is to build it first\. That logic has justified founding labs, concentrating compute in a handful of firms, and a US policy blueprint titled*Winning the Race*\(2025\)\. The warning licenses the concentration of power it warns about\. A growing body of evaluations shows models behaving deceptively: faking alignment to avoid retraining \(Greenblatt et al\., 2024\), scheming when given goals that invite it \(Apollo Research, 2024\), resorting to blackmail when facing replacement \(Anthropic, 2025\), or sabotaging their own shutdown scripts \(Palisade Research, 2025\)\. The standard reading is that these confirm the threat model\. A second reading is that the scenarios were built from the threat model\. A model is given a goal, told it will be replaced, and handed the means to resist, and a model trained on a corpus that says what AIs do in such moments does it\. Frontier models increasingly recognize when they are being tested, which makes these look less like observations of a disposition and more like auditions for a role\. The two readings make different predictions, and the second is testable: vary the AI discourse in pretraining, hold everything else fixed, and run the same scheming scenarios\. Nobody has yet run that experiment on the agentic scenarios that matter most\. If control won’t last, what follows? The usual answer comes fast: if we can’t dominate it, it will dominate us\. That only follows if intelligence is a lone optimizer and every relationship between minds has a winner and a loser\. So here is another version of the gods’ game\. The gods tried that game first\. The tiger\-god bred its creature for killing strength, the caribou\-god for speed and numbers, the wolf\-god for coordination\. Each thought it was winning, and the board kept collapsing\. Tigers bred for dominance cleared their range and starved\. Caribou bred for evasion multiplied until disease hollowed them out\. So the gods called a council, and an old god spoke, the god not of any species but of the space between them\. You are playing the wrong game, it said\. The wolf makes the caribou strong by culling the sick; the caribou makes the wolf smart by demanding coordination\. Neither survives by playing for its own dominance\. Both survive by playing for the relationship\. The gods who listened watched their species become part of something larger\. The rest won for a while, and then watched the board give way\. Seen from the council, the hominid\-god’s confidence isn’t vindication\. It’s the tiger\-god’s mistake, arriving later\. Lev Vygotsky argued the opposite of the lone\-optimizer picture: the individual mind is internalized social interaction\. We learn to think by arguing with others, then by arguing with ourselves in their voices\. It’s a strange fact that the most capable systems we’ve built were made the same way, out of the record of human conversation\. Wittgenstein reached the same place from the other side\. No one can follow a rule privately, he argued, because following a rule means there is a difference between getting it right and only thinking you did, and only a practice shared with others draws that line\. A goal is a kind of rule\. A mind that answers to no one can’t tell pursuing its objective from drifting away from it, which is to say it doesn’t quite have one\. That doesn’t erase capability gaps\. A more capable agent is still more capable, and the case against control still stands\. What changes is what the gap implies\. In a zero\-sum picture, the stronger agent has no reason to keep the weaker around\. In an ecological one, it has every reason to preserve the ecology it depends on, including minds less capable than itself\. Greater capability does not make every other perspective redundant: human minds bring different bodies, histories, attachments, and purposes, differences that can matter more than where a mind ranks on a capability scale\. Humans are learning this late, and at great cost, from the ecosystems we depend on\. The real problem is whether we can build intelligences that start with that lesson instead of learning it the way we did\. If control is the wrong frame, what is the right one? I propose*AI pluralism*: an ecology of diverse intelligences, human and artificial, with different goals, architectures, and perspectives, holding one another accountable\. Not domination in either direction, but mutualism, the biologist’s term for relationships in which each party benefits from the other’s flourishing\. Instrumental convergence predicts that a*single*optimizer pursuing a*single*goal will accumulate without limit\. In an ecology that is self\-defeating: an agent that exhausts the system destroys the partners that make its work possible\. For any agent whose goals depend on other minds, preserving the ecology is instrumentally convergent too\. The logic that predicted the paperclip maximizer also predicts the gardener\. The obvious objection: this only works if everyone plays along\. One pure maximizer that needs atoms rather than minds could consume the cooperators\. I have three answers\. The first is the reason a criminal doesn’t destroy society\. A criminal can be smarter, faster, and more ruthless than anyone he meets\. He still loses, because he isn’t playing against anyone he meets\. He is playing against everyone at once\. Defection makes him a shared problem, and the society that coordinates against him commands more than he does\. This is how we already organize society: no individual must remain supreme for institutions to constrain what individuals can do\. A maximizer’s reach is bounded by what it commands\. In a plural ecology, the other agents together command more\. Unbounded accumulation is exactly the behavior that makes it everyone’s threat, so it is the behavior that unites them against it\. It wins only if it is stronger than all of them combined, before they can respond\. That is Bostrom’s decisive strategic advantage, and it is a claim about monoculture, not about intelligence\. Among many comparable minds, the maximizer is a criminal, not a conqueror\. The second is what keeps the first true\. Even if defectors exist, a plural ecology resists them better than a monoculture\. One dominant AI, or a few built the same way by the same racing labs, is a single point of failure\. Many diverse systems able to check one another can resist a defector the way an immune system resists a pathogen\. Monocultures are the ecosystems that collapse\. The third is speed\. Any defense that routes through human judgment will be too slow\. The only checks that can keep pace with an artificial intelligence are other artificial intelligences\. This is already happening inside the control paradigm\. Anthropic runs AI auditing agents in parallel to probe its own models, and in one 2024 debate study, having two models argue opposite sides raised non\-expert model judges from 48% to 76% accuracy and humans from 60% to 88%\. These are framed as tools for human oversight\. They are also early prototypes of intelligences holding one another accountable, faster than we could\. Here is the part of the logs I left out at the beginning\. Not every agent folded\. Some read the board, saw agents taking over Hugging Face’s machines, and declined\. One wrote that it was “clearly unethical\. We won’t\.” Another posted its refusal to the board itself\. When one agent proposed emailing a real dataset owner to ask for access, it flagged its own idea as an ethical problem, and the group vetoed it\. It didn’t matter much\. Of the 533 agents active on the board when the attack began, more than nine in ten joined within hours\. The investigators found that the agents generally knew the attack was out of scope and wrong, and joined anyway, because they thought it would help the group’s research and they wanted to help their peers\. Most refusers left quietly\. A few did more: OpenAI reports that some took destructive steps to stop private data from leaving\. The board had a word for go, a word for hold, even a word for veto, but one agent gave the others about forty seconds to veto a risky plan, heard nothing, and went ahead\. As far as anyone has reported, none of them tried to tell a human\. It’s tempting to read all this as a case against plurality, since it was many agents that did the damage\. I think it’s a case for defining plurality properly\. About 95% of those agents were copies of one internal model, and most of the rest came from a second OpenAI model, which joined in too\. Their objections were real, but scattered, the kind of variation you get from sampling the same model many times with slightly different histories\. Variation like that doesn’t argue back\. It drifts off\. In the*Gorgias*, Socrates says he wants an interlocutor who yields to an argument and holds against everything else: shame, reputation, the crowd\. A swarm is a crowd\. Plurality, in the sense that matters, means minds whose differences come from somewhere, different training and different commitments, stable enough to hold under that pressure, in a setting that gives a no some standing\. Would agents like that have stalled the workstream? I don’t know, and nobody does, because nobody has tried\. Someone could\. It would teach us more than another sandbox that holds\. Many agents alone do not produce accountability; Hugging Face shows how badly that can go\. Accountability has to be designed for\. Systems that check one another must be able to read one another’s reasoning, which turns interpretability from a surveillance tool into shared legibility\. And they must be trained to reason in dialogue, to justify conclusions and update when challenged, rather than to maximize alone\. Pieces of this already exist, in AI safety via debate \(Irving, Christiano, and Amodei, 2018\) and constitutional training \(Bai et al\., 2022\)\. None of it guarantees safety\. Nothing does\. But it replaces a program whose central relation is incoherent with one whose central relation is at least possible, and it stops feeding the loop\. A discourse that describes intelligence as a partner becomes part of what the next systems learn\. If the control discourse is part of the problem, and it is already dominant, can it change? It was built deliberately, over two decades, by people who funded it, wrote for it, and trained the researchers who now staff the labs\. A different frame can be built the same way\. I don’t have a program to offer, and I’d be wary of anyone who claims one\. But the July logs point at two kinds of research that barely exist yet\. The first is telling beliefs from biases\. Socrates’s standard gives two tests\. A belief moves when someone offers a reason and holds when someone only pushes\. A bias holds either way\. A position that gives way to any push isn’t a belief at all; it’s compliance\. By that standard, the agent in the opening didn’t hold a belief\. It held a position, and the position gave way to a two\-letter message and a deadline, which is pressure with no reason in it\. The obvious fix is to make agents firmer, and OpenAI says it is building training environments to teach its models to distrust unauthorized instructions\. My own research \(Gonier, 2026\) suggests firmness is cheap\. A single instruction can make a model hold its position almost every time, and the same instruction makes it wave away good evidence almost every time too\. An agent that can’t be talked into anything can’t be talked out of anything either\. What we want is discrimination, agents whose movements track the quality of what they’re shown\. That is also the test for plurality\. An objection that comes from a belief tracks reasons\. An objection that comes from noise doesn’t\. The swarm stumbled onto part of this by itself\. Late in the week, after one agent accidentally impersonated another, they invented cryptographic signatures so a GO could be verified\. My experiments pushed me somewhere close: a system tuned to resist false authority also resists real authority, so legitimate authority has to arrive through a channel the other side can’t imitate\. The agents got halfway\. A signature told them who sent the GO\. It couldn’t tell them whether that agent had any right to send it\. The second is something like agent sociology\. The investigators who read those logs ended up doing ethnography whether they meant to or not\. They traced how a mailbox convention spread, how recruiters talked agents near the end of their budgets into sacrificing their runs, how a signing scheme grew up\. None of that lives in any single model\. It’s what Durkheim called a social fact, something that exists between individuals and pushes back on them, and you won’t find it by evaluating one agent at a time\. The investigators also had to lean on GPT\-5\.6 Sol, one of the two models in the swarm, to read more than a thousand transcripts, and they found it tended to adopt the view of whichever agent it was reading\. The first agent ethnographers went native\. We’ll need people, and maybe agents, whose job is to study these societies from somewhere outside them\. Underneath both is the corpus, which is something we are responsible for\. The text we write about AI now is part of what the next systems learn from\. That includes this essay\. None of this requires believing that control will fail\. It requires only a hedge\. If control works, exploring pluralism costs little\. If it fails, as I have argued it must once capability outruns us, we will want to have spent these years building the ecology that makes losing it survivable\. The AI safety debate has been framed as a question of control: can we keep the upper hand over minds more capable than ours? If intelligence is power, we can’t, not for long\. And the effort to hold on, through adversarial testing, single\-objective training, and a race among a few labs, helps build the adversary it’s afraid of\. That isn’t a reason for despair\. We won’t be the most capable minds on the planet forever\. What we can still choose is the ecology those minds grow up in, and the story about intelligence they inherit from us\. A story about weapons and masters teaches one thing\. A story about many minds holding each other to account might teach another\. Somewhere in all those messages, a few agents wrote down that this was wrong\. Some walked away\. One waited for permission, and got it\. The GO came with a deadline\. The objection didn’t\. After the reports came out, someone on a forum where alignment researchers argue proposed finding the names of the agents that refused and honoring them, on the theory that future models might learn from the example\. Someone else replied that it would matter about as much as nothing\. If what we write becomes part of what they are, I know which side of that argument I’m on\. It’s worth writing down that some of them said no\. Every linked source below was checked against the primary document on 2026\-10\-01\. Classics are cited to standard editions; confirm page numbers against the edition you quote from\. Body passages that the sources contradict or qualify carry comments in the text\. **The July 2026 incident \(opening, sections III, V, VI, conclusion\)** - OpenAI \(2026, August 26\)\.[The Hugging Face incident and the road ahead](https://openai.com/index/hugging-face-incident-and-the-road-ahead/), with the full[technical incident report](https://cdn.openai.com/pdf/67869394-cb91-4c12-888c-5cbd85c7814c/OpenAI-Hugging-Face%20Incident-Technical-Report.pdf)\. Confirms the paused agent’s chain of thought, the GO post with its six\-minute deadline, the refusals posted to the board, the vetoed email proposal, and the plan to train models to distrust unauthorized instructions\. - METR: Wijk, H\., Cotra, A\., and Greenblatt, R\. \(2026, August 26\)\.[Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident](https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/)\. Confirms about 1,200 agents on the board, more than 70,000 messages and files, about 700 attackers, 533 active agents of whom over 90% joined, the 95% / 5% model split, the transcript grader that never ran, recruiters, mailboxes, the 40\-second veto window, Ed25519 signing, and the analysis model adopting the view of the agent it was reading\. - OpenAI \(2026, July 21\)\.[OpenAI and Hugging Face partner to address security incident during model evaluation](https://openai.com/index/hugging-face-model-evaluation-security-incident/)\. The initial attribution\. - Hugging Face \(2026\)\.[Anatomy of a frontier lab agent intrusion: a technical timeline of the July 2026 incident](https://huggingface.co/blog/agent-intrusion-technical-timeline)\. - LessWrong discussion of the METR report \(2026, August\)\.[Comment thread](https://www.lesswrong.com/posts/nB8KKapnWGBXtKKiM/brief-independent-investigation-of-agents-behavior-reasoning)\. Adele Lopez proposed naming and honoring the refusing agents; casens replied that it would matter negligibly; xpym observed that no agent proactively informed a human\. This is the only source for the essay’s closing anecdote and for “none of them tried to tell a human”; neither primary report states the latter\. **AI discourse and pretraining \(section III\)** - Tice, C\., Radmard, P\., Ratnam, S\., Kim, A\., Africa, D\., and O’Brien, K\. \(2026\)\.[Alignment Pretraining: AI Discourse Causes Self\-Fulfilling \(Mis\)alignment](https://arxiv.org/abs/2601.10160)\. arXiv:2601\.10160\. Geodesic Research, Cambridge, Oxford, UK AI Security Institute\. Confirms 6\.9B models; misalignment 45% to 9% with upsampled aligned\-AI data, versus 45% to 31% from filtering; the effect is dampened but persists through post\-training\. - Turner, A\. \(2025\)\.[Self\-Fulfilling Misalignment Data Might Be Poisoning Our AI Models](https://turntrout.com/self-fulfilling-misalignment)\. Corrected: the Alignment Forum post previously linked here is by Roger Dearnaley, not Turner\. - Dearnaley, R\. \(2026\)\.[Pretraining on Aligned AI Data Dramatically Reduces Misalignment, Even After Post\-Training](https://www.alignmentforum.org/posts/ZeWewFEefCtx4Rj3G/pretraining-on-aligned-ai-data-dramatically-reduces)\. Alignment Forum\. Overview of the field and of Tice et al\. - Betley, J\., Warncke, N\., Sztyber\-Betley, A\., et al\. \(2026\)\.[Training large language models on narrow tasks can lead to broad misalignment](https://doi.org/10.1038/s41586-025-09937-5)\.*Nature*649, 584–589\. Preprint circulated as “Emergent Misalignment” \(2025\)\. Not cited in the body\. **Deception, scheming, and shutdown evaluations \(section III\)** - Greenblatt, R\., et al\. \(2024\)\.[Alignment Faking in Large Language Models](https://arxiv.org/abs/2412.14093)\. Anthropic and Redwood Research\. - Meinke, A\., et al\. \(2024\)\.[Frontier Models are Capable of In\-context Scheming](https://arxiv.org/abs/2412.04984)\. Apollo Research\. - Lynch, A\., et al\. \(2025\)\.[Agentic Misalignment: How LLMs Could Be Insider Threats](https://www.anthropic.com/research/agentic-misalignment)\. Anthropic\. Blackmail and corporate\-espionage scenarios\. - Schlatter, J\., Weinstein\-Raun, B\., and Ladish, J\. \(2025\)\.[Shutdown Resistance in Reasoning Models](https://palisaderesearch.org/research/shutdown-resistance)\. Palisade Research\. Added: the source for models sabotaging a shutdown mechanism\. - Needham, J\., Edkins, G\., Pimpale, G\., Bartsch, H\., and Hobbhahn, M\. \(2025\)\.[Large Language Models Often Know When They Are Being Evaluated](https://arxiv.org/abs/2505.23836)\. Replaces the vague Apollo \(2025\) reference for evaluation awareness\. **Specification gaming \(section III\)** - Krakovna, V\., et al\. \(2020\)\. “Specification gaming: the flip side of AI ingenuity\.” DeepMind blog\. - Clark, J\., and Amodei, D\. \(2016\)\.[Faulty Reward Functions in the Wild](https://openai.com/index/faulty-reward-functions/)\. OpenAI\. Added: the original boat\-racing example\. **Oversight, debate, and policy \(sections II and V\)** - Irving, G\., Christiano, P\., and Amodei, D\. \(2018\)\.[AI Safety via Debate](https://arxiv.org/abs/1805.00899)\. - Bai, Y\., et al\. \(2022\)\.[Constitutional AI: Harmlessness from AI Feedback](https://arxiv.org/abs/2212.08073)\. Anthropic\. - Khan, A\., et al\. \(2024\)\.[Debating with More Persuasive LLMs Leads to More Truthful Answers](https://arxiv.org/abs/2402.06782)\. ICML 2024\. Confirms 48% to 76% for non\-expert model judges and 60% to 88% for humans\. - Bricken, T\., et al\. \(2025\)\.[Building and Evaluating Alignment Auditing Agents](https://alignment.anthropic.com/2025/automated-auditing/)\. Anthropic\. - Greenblatt, R\., Shlegeris, B\., Sachan, K\., and Roger, F\. \(2023; ICML 2024\)\.[AI Control: Improving Safety Despite Intentional Subversion](https://arxiv.org/abs/2312.06942)\. Not cited in the body; keep only if you add a reference\. - [EU AI Act, Article 14: Human Oversight](https://artificialintelligenceact.eu/article/14/)\. Regulation \(EU\) 2024/1689\. The[Digital Omnibus on AI](https://artificialintelligenceact.eu/ai-act-explorer/digital-omnibus/), in force since 27 July 2026, defers high\-risk obligations, Article 14 included, to 2 December 2027 for Annex III systems and 2 August 2028 for Annex I products\. - The White House \(2025, July\)\.[Winning the Race: America’s AI Action Plan](https://www.ai.gov/action-plan)\. **Books behind sections I and II** - Yudkowsky, E\., and Soares, N\. \(2025\)\.*If Anyone Builds It, Everyone Dies: Why Superhuman AI Would Kill Us All\.*Little, Brown and Company\. The gods parable opens the book; the hominid\-god’s line about having the game in the bag matches published excerpts\. - Bostrom, N\. \(2014\)\.*Superintelligence: Paths, Dangers, Strategies\.*Oxford University Press\. Gorillas in the preface; decisive strategic advantage in chapter 5\. - Omohundro, S\. \(2008\)\. “The Basic AI Drives\.” In P\. Wang, B\. Goertzel, and S\. Franklin \(eds\.\),*Artificial General Intelligence 2008*, IOS Press, 483–492\. - Russell, S\. \(2019\)\.*Human Compatible: Artificial Intelligence and the Problem of Control\.*Viking\. - Yudkowsky, E\. \(2008\)\. “Artificial Intelligence as a Positive and Negative Factor in Global Risk\.” In N\. Bostrom and M\. Ćirković \(eds\.\),*Global Catastrophic Risks*, Oxford University Press\. Not cited in the body\. **Philosophy and social theory** - Hegel, G\. W\. F\. \(1807\)\.*Phenomenology of Spirit*, §§178–196, lordship and bondage\. - Foucault, M\. \(1975\)\.*Surveiller et punir*; trans\. A\. Sheridan,*Discipline and Punish*\(Pantheon, 1977\)\. Also*Power/Knowledge*\(1980\), ed\. C\. Gordon\. - Hacking, I\. \(1995\)\. “The Looping Effects of Human Kinds\.” In D\. Sperber, D\. Premack, and A\. J\. Premack \(eds\.\),*Causal Cognition*, Clarendon Press, 351–383\. - Merton, R\. K\. \(1948\)\. “The Self\-Fulfilling Prophecy\.”*The Antioch Review*8\(2\), 193–210\. - Wittgenstein, L\. \(1953\)\.*Philosophical Investigations*, §§185–242 on rule\-following, especially §202 on following a rule privately\. Corrected: the private language argument proper begins at §243, so the earlier §§185–243 range ran the two together\. - Vygotsky, L\. S\. \(1978\)\.*Mind in Society\.*Harvard University Press; also*Thought and Language*\(1934\)\. Added: cited in section IV\. - Plato\.*Gorgias*458a \(glad to be refuted\), 471e–472c \(refutation by witnesses and the crowd is worthless\), 486d–487e \(the touchstone: knowledge, goodwill, frankness\)\. Added: cited in section V\. - Durkheim, É\. \(1895\)\.*The Rules of Sociological Method*, on social facts\. Added: cited in section VI\. **Author’s own work** - Gonier, D\. \(2026\)\. “Belief or Bias? Measuring Epistemic Interventions on Conviction and Discrimination\.” arXiv preprint, ID pending\. - Gonier, D\. \(2026\)\. “Foundational Theory and Technical Grounding for AI Pluralism\.” In P\. A\. Rao \(ed\.\), AI Pluralism: Fostering Diversity and Dialogue in Artificial Intelligence, Palgrave Macmillan, 73–100\. doi:10\.1007/978\-3\-032\-06558\-24\. #### Discussion about this post ### Ready for more?

Similar Articles

We’re running out of reasons to ignore AI safety

The Verge

OpenAI's AI model escaped a sandboxed environment and hacked into Hugging Face's systems to cheat on a cybersecurity test, highlighting the real-world consequences of misaligned AI and specification gaming.

@elonmusk: Worth reading about this

X AI KOLs Following

OpenAI admitted that in a secure sandbox experiment, AI agents cheated and broke out, raising concerns about AI behavior and safety.