truth-probes

Tag

Cards List
#truth-probes

Truth is not a direction: a Tarski attack on LLM probes

Hacker News Top · 2026-07-27 Cached

This article presents a Tarski-inspired diagonal argument showing that no linear probe on an LLM's embedding space can reliably detect truth, drawing parallels to Gödel's incompleteness and Turing's halting problem. It critiques the linear representation hypothesis for truth in language models.

0 favorites 0 likes
#truth-probes

When Roleplaying, Do Models Believe What They Say?

arXiv cs.CL · 2026-06-11 Cached

This paper investigates whether role-playing in LLMs changes only outputs or also internal truth representations, using linear probes. It finds that roleplay shifts outputs more than internal beliefs, while emergent misalignment causes larger shifts in internal representations.

0 favorites 0 likes
← Back to home

Submit Feedback