Superalignment Fast Grants
Summary
OpenAI announces Superalignment Fast Grants to fund research on aligning superintelligent AI systems, addressing the fundamental challenge of how humans can steer and trust AI systems more capable than themselves. The initiative seeks to rally top researchers to tackle this critical technical problem, which OpenAI believes superintelligence could pose within the next decade.
View Cached Full Text
Cached at: 04/20/26, 02:54 PM
Similar Articles
Weak-to-strong generalization
OpenAI's Superalignment team introduces weak-to-strong generalization, a new research direction for empirically aligning superhuman AI models by addressing the fundamental challenge of how weak human supervisors can reliably control and steer AI systems vastly smarter than themselves.
Advancing independent research on AI alignment
OpenAI is contributing $7.5 million to The Alignment Project, a global independent alignment research fund created by the UK AI Security Institute, helping make it one of the largest dedicated funding efforts for independent alignment research to date. The total fund exceeds £27 million and will support a broad portfolio of alignment research projects worldwide.
Governance of superintelligence
OpenAI outlines a framework for superintelligence governance emphasizing three key pillars: coordination among leading AI development efforts, an international authority (akin to the IAEA) to oversee systems above certain capability thresholds, and technical progress on AI safety with democratic public oversight of the most powerful systems.
Announcing the OpenAI Safety Fellowship
OpenAI announces a new Safety Fellowship program for external researchers to conduct rigorous safety and alignment research on advanced AI systems, running September 2026 through February 2027. The program offers mentorship, compute support, stipends, and workspace at Constellation in Berkeley, with applications open until May 3.
We have 3 years to solve alignment before superintelligence
Geoffrey Irving discusses why short training horizons don't guarantee AI alignment, arguing that long-term plans can be decomposed into short-term subtasks, making instrumental convergence a real concern.