Tag
This article presents a Tarski-inspired diagonal argument showing that no linear probe on an LLM's embedding space can reliably detect truth, drawing parallels to Gödel's incompleteness and Turing's halting problem. It critiques the linear representation hypothesis for truth in language models.
This paper investigates whether role-playing in LLMs changes only outputs or also internal truth representations, using linear probes. It finds that roleplay shifts outputs more than internal beliefs, while emergent misalignment causes larger shifts in internal representations.