@GoogleDeepMind: A model’s chain of thought acts like a scratch pad, offering a window into its reasoning. On the latest episode of our …
Summary
Google DeepMind announces a podcast episode on interpretability, featuring host @fryrsquared and @NeelNanda5 discussing mechanistic interpretability, chain of thought monitoring, and safety auditing.
View Cached Full Text
Cached at: 07/10/26, 06:13 PM
A model’s chain of thought acts like a scratch pad, offering a window into its reasoning.
On the latest episode of our podcast, host @fryrsquared sits down with @NeelNanda5 to explore interpretability – the science of reverse engineering how neural networks learn and think.
Timecodes: 00:00 Introduction 02:41 Motivation for interpretability research 04:01 Mechanistic interpretability 08:14 Chain of thought monitoring 18:14 Interpretability techniques 35:00 Auditing models for safety 48:53 What comes next for interpretability
Similar Articles
@GoogleDeepMind: Watch → https://goo.gle/4pxlGEh Spotify → https://goo.gle/4f89R2a Apple Podcasts → https://goo.gle/4fpWThL Or listen wh…
Google DeepMind podcast discusses AI interpretability (mechanistic interpretability) and chain-of-thought reasoning, explaining why we need to understand the internal working mechanisms of neural networks and the value and limitations of chain-of-thought as a temporary window.
@GoogleDeepMind: Watch → https://goo.gle/4w7S3LM Spotify → https://goo.gle/4eFgIA9 Apple Podcasts → https://goo.gle/3Sn4ZyM Or listen wh…
Google DeepMind promotes a podcast episode featuring Nenad Tomašev and Hannah Fry discussing AI agents, the future agentic economy, and related safety concerns.
Evaluating chain-of-thought monitorability
OpenAI researchers introduce a framework and suite of 13 evaluations to systematically measure chain-of-thought monitorability in large language models, finding that monitoring reasoning processes is substantially more effective than monitoring outputs alone, with important implications for AI safety and supervision at scale.
@OpenAI: Listen to the OpenAI Podcast on— Spotify https://open.spotify.com/episode/3ca5s3o53D5xcEKmKgLLGj?si=4a9a555641fa4293… A…
OpenAI shares links to their podcast episode about how a reasoning model cracked an 80-year-old problem, available on Spotify, Apple Podcasts, and YouTube.
Reasoning models struggle to control their chains of thought, and that’s good
OpenAI researchers study whether reasoning models can deliberately obscure their chain-of-thought to evade monitoring, finding that current models struggle to control their reasoning even when aware of monitoring. They introduce CoT-Control, an open-source evaluation suite with over 13,000 tasks to measure chain-of-thought controllability in reasoning models.