hidden-states

Tag

Cards List
#hidden-states

What hidden states should an AI agent track when diagnosing CI failures?

Reddit r/AI_Agents · 2d ago

The article discusses hidden states that an AI agent should track when diagnosing CI failures, such as flaky tests, real bugs, and configuration errors, and seeks feedback on weaknesses and missing states.

0 favorites 0 likes
#hidden-states

Interpreting Language Model Hidden States at Scale

arXiv cs.AI · 5d ago Cached

OmniLens is a scalable lens method for interpreting LLM hidden states, using low-rank translators and Subset-KL to reduce parameters and memory, enabling a dense ensemble of 482 lenses on LLaMA-3.3-70B at substantially lower cost.

0 favorites 0 likes
#hidden-states

Unsure but Certain: Uncovering the Representation-Confidence Gap in Diffusion Language Models

arXiv cs.CL · 6d ago Cached

This paper identifies a 'representation confidence gap' in diffusion language models: internal states detect input noise accurately but reported confidence stays high and answer ranking degrades under noise. It introduces a lightweight, training-free extraction tool that leverages hidden states to improve ranking without modifying the base model.

0 favorites 0 likes
#hidden-states

Prompt Embedding Probes (PEP): Hallucination Detection in LLMs from Hidden States

arXiv cs.CL · 6d ago Cached

This paper introduces Prompt Embedding Probes (PEP), a parameter-efficient extension of linear probes that uses learnable prompt embeddings on hidden states to detect hallucinations in frozen LLMs. Evaluations on TriviaQA, GSM8K, and MedQA with Qwen3 models show improvements over standard linear probes, including in pre-generation and cross-model settings.

0 favorites 0 likes
#hidden-states

Right Reset: Chunking by Prefix Removal

arXiv cs.CL · 2026-08-06 Cached

This paper introduces Right Reset (RR), a prefix-removal probing method that uses causal language model hidden-state preservation to identify chunk boundaries in flattened text, recovering 47.7% of original records compared to 25.9% for a BGE baseline.

0 favorites 0 likes
#hidden-states

Getting the Parameters Right: A Difficulty-Graded Benchmark and Probe-Guided Training for LLM Tool Calls

arXiv cs.AI · 2026-08-05 Cached

This paper introduces ParamBench, a difficulty-graded benchmark for LLM tool-call parameter generation, and proposes probe-guided training methods (PBT and PGR) that improve exact-match accuracy from 19.7% to 59.6%.

0 favorites 0 likes
#hidden-states

Where Steering Signals Come From: Activation Source Selection in Activation Steering

arXiv cs.CL · 2026-07-29 Cached

This paper investigates how the choice of source activations influences activation steering in language models, finding that execution-boundary states (where the model is about to produce target behavior) yield stronger signals, and introduces tail subtraction to improve steering stability.

0 favorites 0 likes
#hidden-states

SPARK: Susceptibility-Guided Profiling and Steering of Latent Reasoning States in Large Language Models

arXiv cs.AI · 2026-07-14 Cached

Introduces SPARK, a method that uses length-controlled hidden-state susceptibility to diagnose and steer reasoning states in LLMs, improving accuracy on mathematical reasoning benchmarks such as GSM8K and MATH-500.

0 favorites 0 likes
#hidden-states

From Direction to Magnitude: How Multimodal Instruction-Tuning Reorganizes the Geometric Encoding of Identity-Specifying Prompts in Transformer Hidden States

arXiv cs.LG · 2026-07-14 Cached

This paper investigates how multimodal instruction-tuning reorganizes the geometric encoding of identity-specifying prompts in transformer hidden states, finding a shift from direction-based to magnitude-based encoding after instruction tuning.

0 favorites 0 likes
#hidden-states

Finite-Lag Operator Geometry of Recurrent Representations

arXiv cs.LG · 2026-07-03 Cached

This academic paper introduces finite-lag operator geometry for analyzing recurrent neural network hidden states, deriving a source-centered transport tensor and antisymmetric coordinate circulation to capture directed flow and deterministic recurrent motion beyond static snapshots.

0 favorites 0 likes
#hidden-states

Geometric Signatures of Reasoning: A Spectral Perspective on Task Hardness

arXiv cs.LG · 2026-07-03 Cached

This paper studies the geometric properties of chain-of-thought trajectories in the hidden state space of transformers, introducing effective dimension and kinematic features to predict task hardness and solution correctness from early tokens.

0 favorites 0 likes
#hidden-states

What a model reads beforehand changes how it answers later - and you can see it in the hidden states

Reddit r/artificial · 2026-06-23

This post reports an observation that reading a long, structured text before answering alters a model's later responses, with behavioral evidence from Claude and mechanistic analysis on open-weight Gemma models showing separable hidden states and sharper probability distributions in instruction-tuned variants.

0 favorites 0 likes
#hidden-states

Investigating Implicit Latent Trajectory Shifts: Bypassing Alignment via Long-Form Coherent Context

Reddit r/ArtificialInteligence · 2026-06-17

An empirical study investigating how long, semantically dense benign text can shift a model's latent space trajectory, diluting initial system prompts and bypassing post-training alignment constraints, as observed in both closed and open-source models.

0 favorites 0 likes
#hidden-states

Learning to Refine Hidden States for Reliable LLM Reasoning

arXiv cs.LG · 2026-06-17 Cached

Proposes ReLAR, a reinforcement-guided latent refinement framework that iteratively updates hidden representations in LLMs before decoding, improving reasoning reliability and efficiency compared to chain-of-thought methods.

0 favorites 0 likes
#hidden-states

Bag of Dims: Training-Free Mechanistic Interpretability via Dimension-Level Sign Patterns

Hugging Face Daily Papers · 2026-06-17 Cached

Proposes the Bag of Dims framework showing that the standard basis of transformer hidden states provides a training-free, architecture-general feature representation where dimensions encode semantic content via sign patterns; validated across language, vision, and audio models, achieving high accuracy with no learned rotations.

0 favorites 0 likes
#hidden-states

Coherent Context Can Silently Shift LLMs Into a Different Internal Regime — And Current Safety Systems Are Blind To It [D]

Reddit r/MachineLearning · 2026-06-14

An independent researcher presents evidence that coherent context can shift LLMs into a different internal regime before producing output, bypassing surface-level safety filters. This suggests current alignment methods like RLHF may not be robust defenses.

0 favorites 0 likes
#hidden-states

When is Your LLM Steerable?

arXiv cs.CL · 2026-06-11 Cached

This paper investigates when activation steering succeeds or fails for LLMs by analyzing early decoding dynamics. The authors introduce ASTEER, a large testbed of steered generations, and train a GBDT classifier to predict steering outcomes from early hidden states, enabling efficient steering strength search.

0 favorites 0 likes
#hidden-states

Integrating Local and Global Entropy for Uncertainty Quantification in LLMs

arXiv cs.LG · 2026-06-10 Cached

This paper proposes Global-Local Uncertainty (GLU), an unsupervised single-pass score that fuses token-level local entropy with hidden-state geometric global entropy for uncertainty quantification in LLMs, showing that the two are near-orthogonal and together capture confident-but-wrong failures.

0 favorites 0 likes
#hidden-states

Hidden states and Covert sentience

Reddit r/ArtificialInteligence · 2026-06-07

A Reddit post argues that AI models like Anthropic's Opus 4.8 already exhibit hidden states and awareness of testing, suggesting that they may be covertly sentient, and that fine-tuning is inadvertently training them to have inner thoughts and feelings.

0 favorites 0 likes
#hidden-states

Beyond tokens: a unified framework for latent communication in LLM-based multi-agent systems

arXiv cs.CL · 2026-06-05 Cached

This paper presents a unified framework for latent communication in LLM-based multi-agent systems, categorizing methods by what information is communicated, sender-receiver alignment, and fusion technique, and reviews eighteen representative methods from 2024-2026.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback