Tag
This arXiv paper investigates whether simple linear transformations can translate representations across nine heterogeneous text embedding models, finding that shared structure and transferability depend jointly on architecture, training objective, pooling, and data distribution, challenging the notion of universal latent compatibility.
This paper introduces Rift, a method that uses the residual rank of hidden states to detect deceptive responses in language models. It achieves perfect separation across various deception types, model families, and languages, and demonstrates cross-family zero-shot transfer without retraining.