Tag
Research demonstrates that a small set of internal coordinates from LLM hidden states can predict future states and enable targeted edits, with prediction error reduced by 69-76% compared to baseline.
An article describing an experiment where the author asked Claude, an AI model built by Opus 5.5, to reveal insights into its internal workings.
This study explores whether large language models have distinct internal representations of pain, finds that they do, and examines the functional consequences through experiments, concluding with implications for AI safety and welfare.
Discusses research showing that language models exhibit internal states carrying traces of uncertainty, strategic distortion, or misplaced compliance, beyond just bad outputs.
This paper presents a factorised study of probe-based uncertainty estimation in LLMs, showing that raw hidden states and attention features perform well in-domain but structured features are more robust under distribution shift, and provides pretrained probes as off-the-shelf baselines.
This paper introduces POISE, a method for stable policy optimization in large reasoning models by estimating baselines using the model's own internal states, reducing computational overhead compared to PPO and GRPO.
This paper challenges the assumption that LLMs can reliably distinguish between hallucinated and factual outputs through internal signals, arguing that internal states primarily reflect knowledge recall rather than truthfulness. The authors propose a taxonomy of hallucinations (associated vs. unassociated) and show that associated hallucinations exhibit hidden-state geometries overlapping with factual outputs, making standard detection methods ineffective.