poisoning-attacks

Tag

Cards List
#poisoning-attacks

Why Backdooring Neural Networks is so Easy?

arXiv cs.LG ↗ · yesterday Cached

This paper derives an exact closed-form analysis showing that feature learning in neural networks makes them more vulnerable to backdoor attacks, with feature-learning regimes requiring only a trigger strength scaling as α ∝ π^{-1/4} versus π^{-1/2} for lazy learners — theoretically explaining why backdooring large networks is surprisingly easy and why linear security audits underestimate the threat.

0 favorites 0 likes
#poisoning-attacks

PRA-RAG: Provably Robust Aggregation in Retrieval-Augmented Generation against Retrieval Corruption

arXiv cs.AI ↗ · 2026-07-02 Cached

PRA-RAG is a provably robust aggregation algorithm for Retrieval-Augmented Generation that defends against poisoning attacks on retrieved texts. It uses geometric structures in the embedding space to identify robust subsets and provides theoretical bounds on attack impact, reducing attack success rate to as low as 1% while maintaining accuracy.

0 favorites 0 likes
#poisoning-attacks

SkillHarm: Lifecycle-Aware Skill-Based Attacks via Automated Construction

Hugging Face Daily Papers ↗ · 2026-06-01 Cached

SkillHarm is a benchmark for evaluating skill-based attacks across the skill-use lifecycle, revealing high vulnerability (up to 86.3% attack success) in current AI agents and introducing automated attack construction via AutoSkillHarm.

0 favorites 0 likes
← Back to home

Submit Feedback