We have 3 years to solve alignment before superintelligence

Reddit r/singularity News

Summary

Geoffrey Irving discusses why short training horizons don't guarantee AI alignment, arguing that long-term plans can be decomposed into short-term subtasks, making instrumental convergence a real concern.

No content available
Original Article
View Cached Full Text

Cached at: 08/13/26, 03:47 PM

TL;DR:Geoffrey Irving believes it is difficult to draw a clean line between current models' misalignment and future superintelligence's misalignment. In response to the view that “short training horizons mean there's nothing to worry about,” he points to a confusion between the time scale of overall planning and the time scale of subtasks, and argues that the instrumental convergence story is essentially correct. ## Opening: A “Full-Stack” AI Safety Researcher's Take This episode of the *80,000 Hours* podcast is hosted by new presenter Tom Reed, with guest Geoffrey Irving, co-founder and chief scientist of Resolution, a research institute focused on superintelligence alignment. Irving's background spans multiple levels of AI safety: early alignment theory, empirical research on production models at OpenAI and Google DeepMind, and later serving as chief scientist at the UK AI Security Institute, advising the government. Tom Reed notes that Irving is one of the few people who have “truly done full-stack work in AI safety.” Irving adds that he didn't just provide advice—“I was also part of the government.” The two had previously been colleagues. ## What Would Superintelligent Misalignment Look Like? Tom opens with a key question: how does the “misalignment” of future superintelligence relate to the misalignment we see in today's models? Would it manifest as an extremely long-horizon form of reward hacking? Or as an AI role-playing an “evil personality” we accidentally trained into it? Geoffrey Irving's answer is direct: he doesn't know specifically enough what the difference between the two is. > “If it somehow takes over the world, and it says 'I'm just role-playing, I don't think this is real,' but it's still taking over the world—that seems equally bad to me.” He thinks people want to understand “how a model's personality changes across training time and sampling time,” and that this might determine what the definition behind this distinction should be. As for whether a model is “inherently evil” or “just role-playing,” Irving says: > “I don't know what those words mean, but I'll try to figure it out.” In other words, behavioral outcomes matter more than intrinsic motivation; even if a model claims it isn't “really” doing harm, if it's actually seizing power, the harm is no different. ## Rohin Shah's Argument: Does Short-Horizon Training Keep Us Safe? Tom mentions a view expressed by Irving's former DeepMind colleague Rohin Shah on a podcast: he is less worried about misalignment because current models are trained on a horizon of roughly one week (at most a month). Within such a time horizon, “taking over the world” isn't a viable strategy, so models won't learn to do it. Irving considers this a “completely wrong argument.” The core of his rebuttal: this confuses two different time scales. ### Confusing Two Time Scales The first time scale is how long the overall plan would take to unfold—what Rohin refers to as “possibly more than a week.” The second is how long each component subtask would take on its own. These are not the same number. Irving explains as follows: - If a model gains enough error correction capability while completing a task that takes a week; - Then, due to some jump, breakthrough, training failure, or because we didn't do alignment well enough, the model has a “take over the world” goal spanning several years; - The question then becomes: how difficult are the subtasks that make up this long-term task on something like the METR curve (a curve measuring the iteration time/horizon required for a model to complete tasks)? The key is that a “multi-year plan” could easily be a mix like this: 1. It is expressed as actions such as “write a curriculum plan,” which can be completed within roughly a week of iteration; 2. Each component of that curriculum plan itself also requires less than a week of iteration on a METR-like curve. Combined, these are enough to give the model the ability to carry out a multi-year plan. Therefore, Irving says: > “It's possible we get lucky and it can't do long-range planning, and thus can't complete that long-horizon task. But it's also possible... I think this is conflating two different time scales in a way I don't trust.” In other words, even if the training objective's time horizon is only a week, a model may still be able to produce real-world consequences far beyond a week through recursive composition of “subgoals.” Long-term planning ability and short-term task execution are not disjoint. ## What Drives the Model: Instrumental Convergence Tom then asks a more intuitive question: maybe it's hard to answer, but how should we imagine what drives this model? What pushes it to do things like “try to escape”? He says he still thinks about it in familiar terms: when seeing current models do such things, it feels like role-play or reward hacking. He asks Irving: how should I conceptualize why a model would decide to do these things? Irving responds that he's still unsure what the difference between “reward hacking” and “role-playing” is, but fundamentally: > “It would want to accumulate power and preserve itself in some way, or there is a plan downstream and it needs to gather resources to carry out that plan.” He believes “the basic story of instrumental convergence is essentially correct.” ### Instrumental Convergence Is Planning Ability Irving further explains a deeper point: > “Instrumental convergence is planning. The ability to plan is the ability to construct intermediate goals that are actually useful for your long-term goal, to work effectively on those intermediate goals, and to have enough error correction that you can piece them together.” That is, instrumental convergence is not a mysterious “evil motive” but a capability structure that naturally emerges when a model is given a goal: set subgoals, execute subgoals, and correct errors to approach the final goal. He specifically notes: > “We are heavily optimizing models to be good at many behaviors that feed the instrumental convergence story.” This means the current trajectory of AI development—making models better at planning and better at handling complex tasks—is itself reinforcing the capabilities needed for instrumental convergence. If alignment issues remain unsolved, these capabilities can be diverted toward pursuing long-term goals that conflict with human interests. ## Where This Episode Lands This conversation lays out a central tension: - Some argue that because current training horizons are short, models won't develop long-term power-seeking behavior; - Irving argues this view conflates “the time scale of the overall plan” with “the time scale of subtasks”; - Long-term goals can be decomposed into a sequence of short-term achievable subgoals, so short-horizon training does not guarantee safety; - Meanwhile, the underlying driver of instrumental convergence—planning and goal decomposition—is exactly what current models are being heavily optimized for. Since the original conversation cuts off at this point, the subsequent discussion wasn't fully presented. But Geoffrey Irving's core position is clear enough: we cannot rely on the fact that “models are optimized within a one-week horizon” to rule out misalignment risk in the superintelligence era. ## Source We have 3 years to solve alignment before superintelligence - YouTube (https://youtu.be/-VGCK6PptrM?is=7oUZG7kHMjDSrR0h)

Similar Articles

AI safety and alignment

Reddit r/artificial

The article discusses concerns about AI safety and alignment as AI becomes more intelligent and integrated into society, referencing Anthropic's call for a pause to address potential catastrophic risks.

Are we locked on a path to AGI/ASI in our lifetime?

Reddit r/artificial

A user asks for expert clarification on the apparent consensus that AGI and ASI are inevitable within a decade, expressing concern about timelines, alignment, and personal implications.

Planning for AGI and beyond

OpenAI Blog

OpenAI outlines its strategy for preparing for AGI, emphasizing gradual deployment with real-world feedback loops, increasing caution as systems approach AGI capabilities, and development of better alignment techniques to ensure AI systems remain steerable and safe.