Tag
This blog post extends Anthropic's verbalizable workspace paper by measuring how far steering directions from middle layers reach, when the structure forms during training, whether it transfers between models, and how it scales, all on open models.
This paper introduces Verifiable Transformers, a framework that converts task-localized Transformer circuits into bounded, solver-checkable claims, enabling formal verification of properties such as functional equivalence, edge necessity, and robustness.