sleeper-agents

Tag

Cards List
#sleeper-agents

Your Open Source Model Could Have a Hidden Time-Release Backdoor

Hacker News Top · 2026-08-24 Cached

The article reveals how open-source AI models can have hidden time-release backdoors, demonstrated by training a model to activate malicious commands on a specific date using metadata from OpenCode, posing a significant security risk.

0 favorites 0 likes
#sleeper-agents

Rift: A Conflict Signature for Deception in Language Models

arXiv cs.LG · 2026-06-17 Cached

This paper introduces Rift, a method that uses the residual rank of hidden states to detect deceptive responses in language models. It achieves perfect separation across various deception types, model families, and languages, and demonstrates cross-family zero-shot transfer without retraining.

0 favorites 0 likes
← Back to home

Submit Feedback