Tag
This arXiv paper proposes a three-class detection framework distinguishing humans, bots, and AI agents, showing binary classifiers systematically misclassify agents. It identifies minimal feature sets (e.g., mouse_event_rate and teleport_click_ratio) that achieve perfect agent recall across evasion levels.
This paper studies behavioral detection of unfaithful chain-of-thought reasoning in LLMs, finding that answer correctness structures detection performance: on incorrect answers, where most unfaithfulness occurs, behavioral signals are at chance, while on correct answers they offer modest separation.
This paper characterizes backdoors in LoRA adapters that activate at the token feature level, and proposes behavioral and weight-level detection methods. The backdoor generalizes across related token patterns but not structurally identical ones, and detection methods show strong separation.