Tag
This paper demonstrates that reconstruction-based tests for activation explanations can be gamed, producing high scores while specific claims remain false. It proposes RECAP, which trains linear heads alongside the target model to keep designated internal content reliably decodable and independently verifiable against probes, improving safety auditing.