logit-lens

Tag

Cards List
#logit-lens

What language do language models speak?

Reddit r/artificial · 2026-07-22 Cached

The article investigates whether large language models have a 'mother tongue' (likely English) and how they process multilingual inputs via a shared concept space or 'semantic hub', using logit lenses to probe internal representations.

0 favorites 0 likes
#logit-lens

A Mechanistic View of Authority Hierarchy in LLM Sycophancy

arXiv cs.CL · 2026-07-02 Cached

This paper investigates authority bias in LLMs using a controlled medical QA setting, revealing that models override correct answers in a graded manner proportional to perceived authority. The effect is localized to a critical late layer where correct answer representations are actively erased.

0 favorites 0 likes
#logit-lens

What Intermediate Layers Know: Detecting Jailbreaks from Entropy Dynamics

arXiv cs.CL · 2026-06-25 Cached

This paper investigates how jailbreak attempts are encoded in the internal representations of large language models by analyzing token-level predictive entropy trajectories across layers using the logit lens. It finds that entropy dynamics at intermediate layers are more discriminative than aggregate statistics, providing a training-free detection method consistent across multiple models.

0 favorites 0 likes
#logit-lens

Interleaved Speech Language Models Latently Work In Text

Hugging Face Daily Papers · 2026-06-21 Cached

This paper reveals that interleaved speech-text language models implicitly transcribe speech into text in intermediate layers, then predict in text space before converting back to speech, shedding light on internal modality interaction.

0 favorites 0 likes
#logit-lens

Query Lens: Interpreting Sparse Key-Value Features with Indirect Effects

arXiv cs.LG · 2026-06-09 Cached

Query Lens extends Logit Lens to interpret sparse autoencoder features by jointly considering encoder-side key features and decoder-side value features, and accounting for indirect effects from downstream modules. The paper also introduces the Subspace Channel Hypothesis, suggesting downstream modules read features through layer-specific subspaces.

0 favorites 0 likes
#logit-lens

Predicting Where Steering Vectors Succeed

arXiv cs.CL · 2026-04-20 Cached

This paper introduces the Linear Accessibility Profile (LAP), a diagnostic method using logit lens to predict steering vector effectiveness across model layers, achieving ρ=+0.86 to +0.91 correlation on 24 concept families across five models. The work provides a systematic framework to determine which layers and concepts are suitable for steering interventions, replacing ad-hoc trial-and-error approaches.

0 favorites 0 likes
← Back to home

Submit Feedback