data-poisoning

Tag

Cards List
#data-poisoning

On the Value of Human Ideas: What data poisoning research reveals about "autonomous" AI breakthroughs

Reddit r/LocalLLaMA · 2026-09-08

The article discusses how data poisoning research reveals that small amounts of targeted data can disproportionately influence AI models, suggesting that accumulated human ideas from user interactions might contribute to AI breakthroughs, challenging the notion of data dilution.

0 favorites 0 likes
#data-poisoning

Where do yall stop poisoning - at the agent, the perimeter, ?

Reddit r/AI_Agents · 2026-08-13

A hypothetical security discussion about how poisoned data from a compromised API could flow into AI agents and analytics systems, questioning where defenses should be placed.

0 favorites 0 likes
#data-poisoning

The web’s newest weapon against AI scrapers is a font

Ars Technica · 2026-08-12 Cached

A new font called ShieldFont leverages ligatures to replace words in webpage HTML, presenting readable text to humans while serving nonsense to AI scrapers, effectively poisoning training data.

0 favorites 0 likes
#data-poisoning

Data poisoning and RAG manipulation

Reddit r/ArtificialInteligence · 2026-08-09

A discussion of how data poisoning and RAG manipulation pose a silent, dangerous threat to AI systems, arguing that security must extend beyond input filtering to memory, data pipelines, and multi-agent logic.

0 favorites 0 likes
#data-poisoning

It’s Time to Poison AI | GN Mega Charts Update & LLM Countermeasures

Reddit r/ArtificialInteligence · 2026-08-05 Cached

Gamers Nexus announces it is starting to poison its free benchmark charts to deter AI scrapers, with subtle pixel-level data manipulations tested against LLMs, as bot traffic now exceeds human traffic online.

0 favorites 0 likes
#data-poisoning

RAGuard: A Layered Defense Framework for Retrieval-Augmented Generation Systems Against Data Poisoning

arXiv cs.LG · 2026-07-30 Cached

RAGuard is a layered defense framework for Retrieval-Augmented Generation (RAG) systems that uses adversarial fine-tuning of the retriever and a label-free filter (ZKIP) to achieve zero attack success against corpus poisoning, maintaining high retrieval accuracy.

0 favorites 0 likes
#data-poisoning

AI Companies Are Buying Tons of Old Books Because They're Free of AI Slop

Reddit r/ArtificialInteligence · 2026-07-21 Cached

AI companies are buying pre-2022 printed books to avoid AI-generated text in training data, as old books are guaranteed free of AI slop and poisoning. ISBNdb offers bulk book acquisition services to AI labs under NDAs.

0 favorites 0 likes
#data-poisoning

Two Sides of the Same Coin: Learning the Backdoor to Remove the Backdoor

arXiv cs.LG · 2026-07-08 Cached

HARVEY learns a backdoored reference model to accurately identify poisonous samples, achieving near-perfect backdoor removal with minimal accuracy loss.

0 favorites 0 likes
#data-poisoning

Mental Damage: Caption Poisoning Attacks on Retrieval-Augmented Text-to-Music Generation

arXiv cs.AI · 2026-06-01 Cached

This paper introduces a dual-layer caption poisoning attack on retrieval-augmented text-to-music systems, showing that an attacker can inject malicious captions into the knowledge database to steer generated music toward attacker-chosen intent without modifying user prompts or models.

0 favorites 0 likes
#data-poisoning

Be Kind, Rewrite: Benign Projections via Rewriting Defend Against LLM Data Poisoning Attacks

Hugging Face Daily Papers · 2026-05-18 Cached

This paper introduces open-book benign rewriting (OBBR) as a proactive defense against backdoor attacks on LLMs, showing it neutralizes harmful content by projecting to benign prompts, and improves safety by 51% over state-of-the-art defenses.

0 favorites 0 likes
#data-poisoning

What are AI tarpits? Understanding the tools people are using to poison LLMs

Reddit r/ArtificialInteligence · 2026-05-17 Cached

AI tarpits are tools used by content creators to poison large language models by feeding scrapers useless or incorrect data, degrading AI output quality.

0 favorites 0 likes
#data-poisoning

SoK: A Comprehensive Analysis of the Current Status of Neural Tangent Generalization Attacks with Research Directions

arXiv cs.LG · 2026-05-14 Cached

This paper presents a comprehensive analysis of the Neural Tangent Generalization Attack (NTGA) for data protection, including a taxonomy of related attacks, and discusses future research directions.

0 favorites 0 likes
#data-poisoning

When Emotion Becomes Trigger: Emotion-style dynamic Backdoor Attack Parasitising Large Language Models

arXiv cs.CL · 2026-05-13 Cached

This paper introduces Paraesthesia, a dynamic backdoor attack on LLMs that uses emotional style as a stealthy trigger during fine-tuning, achieving high success rates while maintaining model utility.

0 favorites 0 likes
← Back to home

Submit Feedback