Tag
This paper proposes a framework to analyze how linguistic relations are linearly encoded in language model embeddings, revealing differences across models like GloVe, RoBERTa, and ModernBERT and relation types.
The paper shows that the hallucination signal in LLMs is dominated by a mean shift, making simple probes like L2-regularized logistic regression sufficient and outperforming complex alternatives.