Tag
This paper investigates how language models separate a character's belief from reality, finding that they use a shared value slot for attributed values and a router at the query position to select the frame (belief or reality) to read out. It identifies two routes for asserted and derived beliefs, and shows that the slot itself carries no belief-reality tag; the separation lies in dissociated routing subspaces.
This paper investigates how transformer language models implement multiple mental spaces (counterfactual, belief, fictional, etc.) using a shared router mechanism over a value slot, showing that a single low-rank subspace controls all space types and drives inference.