hidden-states

Tag

Cards List
#hidden-states

Rethinking Reasoning Paths as Phase-Structured Trajectories

arXiv cs.AI ↗ · 3d ago Cached

The paper proposes PAIR (Phase-Aligned Intra-question Reasoning), which treats LLM reasoning paths as phase-structured trajectories and evaluates path quality only within the same question, showing that standard correctness probes partly rely on question-level variation and that phase-wise steering yields causal evidence of path-quality directions.

0 favorites 0 likes
#hidden-states

@rohanpaul_ai: You can predict where an LLM's internal state is heading, and that targeted edits can pull it back on course. On a froz…

X AI KOLs Following ↗ · 3d ago Cached

Research demonstrates that a small set of internal coordinates from LLM hidden states can predict future states and enable targeted edits, with prediction error reduced by 69-76% compared to baseline.

0 favorites 0 likes
#hidden-states

Does the Truthfulness Signal Survive Code-Mixing? Probing Hidden States for Hallucination Detection in Hinglish

arXiv cs.CL ↗ · 2026-09-22 Cached

The study investigates whether hallucination detection probes trained on monolingual text transfer to Hinglish, finding robust performance and higher hallucination rates in Hindi and Hinglish compared to English.

0 favorites 0 likes
#hidden-states

Bypass Observation: A Conceptual Design of a Non-Intrusive Layer-Wise Semantic Extraction Architecture

arXiv cs.AI ↗ · 2026-09-15 Cached

This paper proposes Bypass Observation, a non-intrusive architecture for extracting intermediate semantic information from large language models to improve interpretability and safety auditing.

0 favorites 0 likes
#hidden-states

ForeSight: Enhancing Risk Monitoring via Early Safety Signal Distillation

arXiv cs.CL ↗ · 2026-09-15 Cached

ForeSight is a framework that predicts harmful outputs from large language models by analyzing first-token hidden states, improving early risk monitoring efficiency.

0 favorites 0 likes
#hidden-states

Reference-Based Bias Detection in LLMs via Relative Representations of Hidden States

Hugging Face Daily Papers ↗ · 2026-09-09 Cached

The paper proposes a reference-based hidden-state auditing method to detect bias shifts in LLMs with minimal compute, showing high correlation with output-level bias changes across benchmarks.

0 favorites 0 likes
#hidden-states

What hidden states should an AI agent track when diagnosing CI failures?

Reddit r/AI_Agents ↗ · 2026-08-14

The article discusses hidden states that an AI agent should track when diagnosing CI failures, such as flaky tests, real bugs, and configuration errors, and seeks feedback on weaknesses and missing states.

0 favorites 0 likes
#hidden-states

Interpreting Language Model Hidden States at Scale

arXiv cs.AI ↗ · 2026-08-12 Cached

OmniLens is a scalable lens method for interpreting LLM hidden states, using low-rank translators and Subset-KL to reduce parameters and memory, enabling a dense ensemble of 482 lenses on LLaMA-3.3-70B at substantially lower cost.

0 favorites 0 likes
#hidden-states

Unsure but Certain: Uncovering the Representation-Confidence Gap in Diffusion Language Models

arXiv cs.CL ↗ · 2026-08-11 Cached

This paper identifies a 'representation confidence gap' in diffusion language models: internal states detect input noise accurately but reported confidence stays high and answer ranking degrades under noise. It introduces a lightweight, training-free extraction tool that leverages hidden states to improve ranking without modifying the base model.

0 favorites 0 likes
#hidden-states

Prompt Embedding Probes (PEP): Hallucination Detection in LLMs from Hidden States

arXiv cs.CL ↗ · 2026-08-11 Cached

This paper introduces Prompt Embedding Probes (PEP), a parameter-efficient extension of linear probes that uses learnable prompt embeddings on hidden states to detect hallucinations in frozen LLMs. Evaluations on TriviaQA, GSM8K, and MedQA with Qwen3 models show improvements over standard linear probes, including in pre-generation and cross-model settings.

0 favorites 0 likes
#hidden-states

Right Reset: Chunking by Prefix Removal

arXiv cs.CL ↗ · 2026-08-06 Cached

This paper introduces Right Reset (RR), a prefix-removal probing method that uses causal language model hidden-state preservation to identify chunk boundaries in flattened text, recovering 47.7% of original records compared to 25.9% for a BGE baseline.

0 favorites 0 likes
#hidden-states

Getting the Parameters Right: A Difficulty-Graded Benchmark and Probe-Guided Training for LLM Tool Calls

arXiv cs.AI ↗ · 2026-08-05 Cached

This paper introduces ParamBench, a difficulty-graded benchmark for LLM tool-call parameter generation, and proposes probe-guided training methods (PBT and PGR) that improve exact-match accuracy from 19.7% to 59.6%.

0 favorites 0 likes
#hidden-states

Where Steering Signals Come From: Activation Source Selection in Activation Steering

arXiv cs.CL ↗ · 2026-07-29 Cached

This paper investigates how the choice of source activations influences activation steering in language models, finding that execution-boundary states (where the model is about to produce target behavior) yield stronger signals, and introduces tail subtraction to improve steering stability.

0 favorites 0 likes
#hidden-states

SPARK: Susceptibility-Guided Profiling and Steering of Latent Reasoning States in Large Language Models

arXiv cs.AI ↗ · 2026-07-14 Cached

Introduces SPARK, a method that uses length-controlled hidden-state susceptibility to diagnose and steer reasoning states in LLMs, improving accuracy on mathematical reasoning benchmarks such as GSM8K and MATH-500.

0 favorites 0 likes
#hidden-states

From Direction to Magnitude: How Multimodal Instruction-Tuning Reorganizes the Geometric Encoding of Identity-Specifying Prompts in Transformer Hidden States

arXiv cs.LG ↗ · 2026-07-14 Cached

This paper investigates how multimodal instruction-tuning reorganizes the geometric encoding of identity-specifying prompts in transformer hidden states, finding a shift from direction-based to magnitude-based encoding after instruction tuning.

0 favorites 0 likes
#hidden-states

Finite-Lag Operator Geometry of Recurrent Representations

arXiv cs.LG ↗ · 2026-07-03 Cached

This academic paper introduces finite-lag operator geometry for analyzing recurrent neural network hidden states, deriving a source-centered transport tensor and antisymmetric coordinate circulation to capture directed flow and deterministic recurrent motion beyond static snapshots.

0 favorites 0 likes
#hidden-states

Geometric Signatures of Reasoning: A Spectral Perspective on Task Hardness

arXiv cs.LG ↗ · 2026-07-03 Cached

This paper studies the geometric properties of chain-of-thought trajectories in the hidden state space of transformers, introducing effective dimension and kinematic features to predict task hardness and solution correctness from early tokens.

0 favorites 0 likes
#hidden-states

What a model reads beforehand changes how it answers later - and you can see it in the hidden states

Reddit r/artificial ↗ · 2026-06-23

This post reports an observation that reading a long, structured text before answering alters a model's later responses, with behavioral evidence from Claude and mechanistic analysis on open-weight Gemma models showing separable hidden states and sharper probability distributions in instruction-tuned variants.

0 favorites 0 likes
#hidden-states

Investigating Implicit Latent Trajectory Shifts: Bypassing Alignment via Long-Form Coherent Context

Reddit r/ArtificialInteligence ↗ · 2026-06-17

An empirical study investigating how long, semantically dense benign text can shift a model's latent space trajectory, diluting initial system prompts and bypassing post-training alignment constraints, as observed in both closed and open-source models.

0 favorites 0 likes
#hidden-states

Learning to Refine Hidden States for Reliable LLM Reasoning

arXiv cs.LG ↗ · 2026-06-17 Cached

Proposes ReLAR, a reinforcement-guided latent refinement framework that iteratively updates hidden representations in LLMs before decoding, improving reasoning reliability and efficiency compared to chain-of-thought methods.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback