@timoreilly: About to go live with @mlpowered to talk about the "neuroscience of AI." https://learning.oreilly.com/live-events/crack…
Summary
Tim O’Reilly and Emmanuel Ameisen discuss AI interpretability, focusing on world models in LLMs like Claude and Anthropic's tools for steering model behavior.
View Cached Full Text
Cached at: 09/09/26, 05:51 PM
About to go live with @mlpowered to talk about the “neuroscience of AI.” https://learning.oreilly.com/live-events/cracking-llms-open-emmanuel-ameisen-live-with-tim-oreilly/0642572444716/… Join for free.
Cracking LLMs Open: Emmanuel Ameisen Live with Tim O’Reilly
Source: https://www.oreilly.com/live-events/cracking-llms-open-emmanuel-ameisen-live-with-tim-oreilly/0642572444716/ View all eventsPublished byO’Reilly Media, Inc.
Intermediate content levelIntermediate
What Claude’s internals reveal about how it actually thinks
ScheduleAnthropic’s interpretability team studies Claude’s internals the way neuroscientists study a brain, but the science is younger, messier, and far less mapped. Emmanuel Ameisen has been part of that effort for the last two years, and at O’Reilly’s recent Foo Camp, he shared some of what he’s learned about what happens inside of Claude when it processes information.
Token prediction is often framed as simply “pattern matching,” Emmanuel noted in his talk, but it turns out that in order to be superhuman at token prediction, models build a complex understanding of the world. Ask Claude to finish a sentence about a short bike ride across the “GG bridge” and it correctly infers that the bridge must be the Golden Gate, placing the user in San Francisco. From there, the model can build on this understanding to answer questions, for instance how long it will take to get to “world-class skiing” at Lake Tahoe. His team can see this world model emerge inside the model as different kinds of activation patterns light up when the model completes complex tasks.
Tim asked Emmanuel to join him for this episode to extend his Foo Camp presentation. They’ll start by establishing some foundations about what we mean when we talk about world models, and whether what we find inside language models matches that description. They’ll discuss examples of these world models at work when doing math, or writing poetry, where the models choose how a line will end before writing it. They’ll also discuss the tools Anthropic’s interpretability team has built to develop this understanding, steering models by “pushing or pulling” on their activations to get them to behave accordingly. Along the way, Emmanuel and Tim will think big-picture about interpretability: What should people take on faith about model behavior right now, and what should they not?
Join in to find out what Emmanuel’s learned from two years of looking inside Claude. And bring your questions and comments. They’ll help guide the conversation.
Recommended prep or follow-up:
Schedule
The time frames are only estimates and may vary according to how the class is progressing.
Wednesday, September 9, 2026, at 9:00am PT / 12:00pm ET
- Interactive discussion and Q&A (60 minutes)
Tim O’Reillyis the founder and CEO of O’Reilly Media, Inc. His original business plan was simply “interesting work for interesting people,” and that’s worked out pretty well. He publishes books, runs online conferences, invests in early-stage startups, urges companies to create more value than they capture, and tries to change the world by spreading and amplifying the knowledge of innovators. He’s perhaps best known for his role in shaping big ideas like open source software, unconferences (Foo Camp), Web 2.0, and government as a platform. His 2017 bookWTF? What’s the Future and Why It’s Up to Usexplored the role of human agency in shaping the future in the face of the coming AI wave. These days, he’s focused on mechanism design for the human-AI economy. Mechanism design is sometimes described as “reverse game theory”, whereby you start with the outcome you want and then figure out what rules of the game will produce that outcome. He explores these ideas at the non-profit AI Disclosures Project, which he co-founded with Ilan Strauss. He writes frequently on Substack at theO’Reilly Radar, the AI Disclosures Project’sAsimov’s Addendum, and his own Conversations with AI.
Emmanuel Ameisen works on mechanistic interpretability at Anthropic: building tools that trace the computations inside large language models, and using them to study how those systems do what they do. His work includes finding evidence that LLMs plan ahead, an account of the general circuits they learn for representing and manipulating numbers, and explanations for some of the surprising behavior we see in frontier systems. Before Anthropic, he was a Staff ML Engineer at Stripe, and led Insight Data Science’s AI program, directing more than a hundred applied ML projects. He is the author of the O’Reilly bookBuilding Machine Learning Powered Applications.
Skill covered
Computer Vision
Similar Articles
@timoreilly: I wrote this post (The Collaborative Exoskeleton of AI Science) a month or so ago and then forgot to publish it! It’s w…
Tim O'Reilly discusses the challenges of integrating AI into scientific publishing, including hallucinated citations, propagation of retracted papers, and training on compromised literature, and calls for adapting existing scientific infrastructure for AI use.
@timodonnell: Want to watch a 535B parameter (23B active) LLM get trained live? Follow along here https://wandb.ai/marin-community/ma…
A tweet announces the live training of a 535B parameter (23B active) large language model, with links to follow the process on Weights & Biases and GitHub.
@timoreilly: I made an offhand comment about AI for writing in my recent interview with @stevenlevy. A number of people have comment…
Tim O'Reilly discusses AI as a medium for writing, comparing it to traditional art forms and defending its creative use despite concerns from professional writers.
@danintheory: Great conversation and a fun way to learn about an important open AI problem!
Sequoia Capital highlights the gap between current AI models that train once and human continuous learning, and points to EngramLab's work on AI that never stops learning with memory inside the model.
@Prince_Canuma: My @aiDotEngineer talk is live: "On-device Intelligence using MLX" Huge thanks to @swyx and the team for having me — ha…
The author announces their live talk titled 'On-device Intelligence using MLX' at the aiDotEngineer event, expressing gratitude to the organizers and community contributors.