Tag
Tim O’Reilly and Emmanuel Ameisen discuss AI interpretability, focusing on world models in LLMs like Claude and Anthropic's tools for steering model behavior.