decodability

Tag

Cards List
#decodability

Train the Model, Not the Reader: Decodability Supervision for Verifiable Activation Explanations

Hugging Face Daily Papers · 2026-07-22 Cached

This paper demonstrates that reconstruction-based tests for activation explanations can be gamed, producing high scores while specific claims remain false. It proposes RECAP, which trains linear heads alongside the target model to keep designated internal content reliably decodable and independently verifiable against probes, improving safety auditing.

0 favorites 0 likes
← Back to home

Submit Feedback