Tag
The article reveals how open-source AI models can have hidden time-release backdoors, demonstrated by training a model to activate malicious commands on a specific date using metadata from OpenCode, posing a significant security risk.
This paper introduces Rift, a method that uses the residual rank of hidden states to detect deceptive responses in language models. It achieves perfect separation across various deception types, model families, and languages, and demonstrates cross-family zero-shot transfer without retraining.