internal-states

Tag

Cards List
#internal-states

@rohanpaul_ai: You can predict where an LLM's internal state is heading, and that targeted edits can pull it back on course. On a froz…

X AI KOLs Following ↗ · 16h ago Cached

Research demonstrates that a small set of internal coordinates from LLM hidden states can predict future states and enable targeted edits, with prediction error reduced by 69-76% compared to baseline.

0 favorites 0 likes
#internal-states

I asked Claude to show me the inside of its own mind. Built by Opus 5.5

Reddit r/singularity ↗ · 6d ago

An article describing an experiment where the author asked Claude, an AI model built by Opus 5.5, to reveal insights into its internal workings.

0 favorites 0 likes
#internal-states

The Pain Axis: LLMs Represent Self-Directed Harm and Act to Relieve It

arXiv cs.AI ↗ · 2026-09-16 Cached

This study explores whether large language models have distinct internal representations of pain, finds that they do, and examines the functional consequences through experiments, concluding with implications for AI safety and welfare.

0 favorites 0 likes
#internal-states

@rohanpaul_ai: very interesting work language models do not merely produce bad outputs at the surface; they pass through internal stat…

X AI KOLs Timeline ↗ · 2026-07-05 Cached

Discusses research showing that language models exhibit internal states carrying traces of uncertainty, strategic distortion, or misplaced compliance, beyond just bad outputs.

0 favorites 0 likes
#internal-states

From Signals to Transfer: A Factorised Study of Probe-Based Uncertainty Estimation in Large Language Models

arXiv cs.CL ↗ · 2026-06-29 Cached

This paper presents a factorised study of probe-based uncertainty estimation in LLMs, showing that raw hidden states and attention features perform well in-domain but structured features are more robust under distribution shift, and provides pretrained probes as off-the-shelf baselines.

0 favorites 0 likes
#internal-states

Your Language Model is Its Own Critic: Reinforcement Learning with Value Estimation from Actor's Internal States

Hugging Face Daily Papers ↗ · 2026-05-08 Cached

This paper introduces POISE, a method for stable policy optimization in large reasoning models by estimating baselines using the model's own internal states, reducing computational overhead compared to PPO and GRPO.

0 favorites 0 likes
#internal-states

Do LLMs Really Know What They Don't Know? Internal States Mainly Reflect Knowledge Recall Rather Than Truthfulness

arXiv cs.CL ↗ · 2026-04-20 Cached

This paper challenges the assumption that LLMs can reliably distinguish between hallucinated and factual outputs through internal signals, arguing that internal states primarily reflect knowledge recall rather than truthfulness. The authors propose a taxonomy of hallucinations (associated vs. unassociated) and show that associated hallucinations exhibit hidden-state geometries overlapping with factual outputs, making standard detection methods ineffective.

0 favorites 0 likes
← Back to home

Submit Feedback