Tag
Introduces AntiSkillBench, a benchmark for evaluating privacy leakage, impersonation risk, and defenses in persona-skill pipelines for AI agents, finding that risks persist across backbones and existing defenses are limited.
This paper investigates OS resilience against self-state attacks on self-hosted AI agents, characterizing an attack space and evaluating layered defense strategies. It finds that while a layered defense stack is effective, a small residual attack surface remains structurally indistinguishable at the OS level.
A discussion about handling prompt injection attacks in AI agents that read external content like emails and webpages, exploring production-level defenses and the subtle threats beyond obvious patterns.
Hacker Vitto Rivabella publicly announced the successful jailbreak of Fable 5, analyzing in detail the model's multi-layered security mechanisms, including input/output auditing, intent detection, and chain-of-thought defense, and provided methods to bypass them.