Tag
The article discusses how data poisoning research reveals that small amounts of targeted data can disproportionately influence AI models, suggesting that accumulated human ideas from user interactions might contribute to AI breakthroughs, challenging the notion of data dilution.
A hypothetical security discussion about how poisoned data from a compromised API could flow into AI agents and analytics systems, questioning where defenses should be placed.
A new font called ShieldFont leverages ligatures to replace words in webpage HTML, presenting readable text to humans while serving nonsense to AI scrapers, effectively poisoning training data.
A discussion of how data poisoning and RAG manipulation pose a silent, dangerous threat to AI systems, arguing that security must extend beyond input filtering to memory, data pipelines, and multi-agent logic.
Gamers Nexus announces it is starting to poison its free benchmark charts to deter AI scrapers, with subtle pixel-level data manipulations tested against LLMs, as bot traffic now exceeds human traffic online.
RAGuard is a layered defense framework for Retrieval-Augmented Generation (RAG) systems that uses adversarial fine-tuning of the retriever and a label-free filter (ZKIP) to achieve zero attack success against corpus poisoning, maintaining high retrieval accuracy.
AI companies are buying pre-2022 printed books to avoid AI-generated text in training data, as old books are guaranteed free of AI slop and poisoning. ISBNdb offers bulk book acquisition services to AI labs under NDAs.
HARVEY learns a backdoored reference model to accurately identify poisonous samples, achieving near-perfect backdoor removal with minimal accuracy loss.
This paper introduces a dual-layer caption poisoning attack on retrieval-augmented text-to-music systems, showing that an attacker can inject malicious captions into the knowledge database to steer generated music toward attacker-chosen intent without modifying user prompts or models.
This paper introduces open-book benign rewriting (OBBR) as a proactive defense against backdoor attacks on LLMs, showing it neutralizes harmful content by projecting to benign prompts, and improves safety by 51% over state-of-the-art defenses.
AI tarpits are tools used by content creators to poison large language models by feeding scrapers useless or incorrect data, degrading AI output quality.
This paper presents a comprehensive analysis of the Neural Tangent Generalization Attack (NTGA) for data protection, including a taxonomy of related attacks, and discusses future research directions.
This paper introduces Paraesthesia, a dynamic backdoor attack on LLMs that uses emotional style as a stealthy trigger during fine-tuning, achieving high success rates while maintaining model utility.