Tag
This paper derives an exact closed-form analysis showing that feature learning in neural networks makes them more vulnerable to backdoor attacks, with feature-learning regimes requiring only a trigger strength scaling as α ∝ π^{-1/4} versus π^{-1/2} for lazy learners — theoretically explaining why backdooring large networks is surprisingly easy and why linear security audits underestimate the threat.
PRA-RAG is a provably robust aggregation algorithm for Retrieval-Augmented Generation that defends against poisoning attacks on retrieved texts. It uses geometric structures in the embedding space to identify robust subsets and provides theoretical bounds on attack impact, reducing attack success rate to as low as 1% while maintaining accuracy.
SkillHarm is a benchmark for evaluating skill-based attacks across the skill-use lifecycle, revealing high vulnerability (up to 86.3% attack success) in current AI agents and introducing automated attack construction via AutoSkillHarm.