defenses

Tag

Cards List
#defenses

When Agents Learn to Be You: Benchmarking Privacy Leakage, Impersonation Risk, and Defenses in Persona Skills

Hugging Face Daily Papers · 2026-08-04 Cached

Introduces AntiSkillBench, a benchmark for evaluating privacy leakage, impersonation risk, and defenses in persona-skill pipelines for AI agents, finding that risks persist across backbones and existing defenses are limited.

0 favorites 0 likes
#defenses

Self-State Attacks on Self-Hosted AI Agents: How Far Can OS Defenses Go?

Hugging Face Daily Papers · 2026-07-20 Cached

This paper investigates OS resilience against self-state attacks on self-hosted AI agents, characterizing an attack space and evaluating layered defense strategies. It finds that while a layered defense stack is effective, a small residual attack surface remains structurally indistinguishable at the OS level.

0 favorites 0 likes
#defenses

How are you all handling prompt injection for agents that read external content?

Reddit r/AI_Agents · 2026-07-04

A discussion about handling prompt injection attacks in AI agents that read external content like emails and webpages, exploring production-level defenses and the subtle threats beyond obvious patterns.

0 favorites 0 likes
#defenses

@MaxForAI: Hacker Vitto Rivabella publicly announced that Fable 5 has been broken again. He said most jailbreak attempts have failed, and the defenses are clearly layered. The model is extremely well protected (of course it blocks 90% of requests, but they did a really good job). The model seems to perform security checks on both input and output...

X AI KOLs Timeline · 2026-07-03 Cached

Hacker Vitto Rivabella publicly announced the successful jailbreak of Fable 5, analyzing in detail the model's multi-layered security mechanisms, including input/output auditing, intent detection, and chain-of-thought defense, and provided methods to bypass them.

0 favorites 0 likes
← Back to home

Submit Feedback