memory-poisoning

Tag

Cards List
#memory-poisoning

CAPTURE: Disentangling Preference Drift from Memory Poisoning in Personalized LLM Agents

arXiv cs.LG · 2026-09-03 Cached

The paper introduces CAPTURE, a method to distinguish between genuine user preference changes and malicious memory poisoning in personalized language agents using neural differential equations and causal auditing.

0 favorites 0 likes
#memory-poisoning

FARMA: a memory-poisoning attack that forges an agent's own decision logs, not its retrieved facts

Reddit r/AI_Agents · 2026-07-11

FARMA is a novel memory-poisoning attack that targets an agent's own decision logs rather than retrieved facts, achieving 100% attack success against undefended and defended systems, with the authors' defense SENTINEL reducing success to 0% but remaining vulnerable to adaptive attackers.

0 favorites 0 likes
#memory-poisoning

When Agents Remember Too Much: Memory Poisoning Attacks on Large Language Model Agents

arXiv cs.AI · 2026-07-09 Cached

This paper introduces GhostWriter, a novel attack vector that exploits memory subsystems in LLM-powered personal agents to poison their memory store, achieving high injection and activation rates. The authors propose AM-Sentry, a defense that significantly reduces attack success while maintaining agent utility.

0 favorites 0 likes
#memory-poisoning

@rohanpaul_ai: A warning for anyone using autonomous agents Google DeepMind’s paper. Gives the first clear taxonomy of 6 attack types …

X AI KOLs Timeline · 2026-07-06 Cached

Google DeepMind's paper provides the first clear taxonomy of six attack types on autonomous AI agents, revealing that harmful websites can hide content like instructions in HTML comments or steganography that agents parse but humans never see, achieving up to 86% agent commandeering in benchmarks.

0 favorites 0 likes
#memory-poisoning

ElephantAgent: Contextual State Continuity in Agentic Systems

arXiv cs.AI · 2026-07-03 Cached

This paper presents ElephantAgent, a protocol that enforces Contextual State Continuity to defend against contextual state poisoning in agentic systems, using replicated trusted hardware and historical traceability.

0 favorites 0 likes
#memory-poisoning

2 papers this week that fix agent memory poisoning & privacy leaks, but there's no library you can actually use.

Reddit r/AI_Agents · 2026-06-27

Two papers published this week address agent memory poisoning and privacy leaks, but no readily usable library or tool has been released to implement these fixes.

0 favorites 0 likes
#memory-poisoning

The Containment Gap: How Deployed Agentic AI Frameworks Fail Public-Facing Safety Requirements

arXiv cs.AI · 2026-06-12 Cached

This paper audits LangChain, AutoGPT, and OpenAI Agents SDK for architectural safety guarantees and finds no native compliance with containment principles, demonstrating that memory poisoning can cause persistent failures; it introduces lightweight mechanisms to eliminate such attacks.

0 favorites 0 likes
#memory-poisoning

The Misattribution Gap: When Memory Poisoning Looks Like Model Failure in Agentic AI Systems

arXiv cs.AI · 2026-05-25 Cached

This paper identifies a structural failure in multi-agent AI pipelines where memory-layer attacks can be misattributed as model misalignment, formalizing Semantic Norm Drift (SND) and proposing Counterfactual Composition Testing and Memory-Persistent Information-Flow Control as defenses.

0 favorites 0 likes
#memory-poisoning

MemAudit: Post-hoc Auditing of Poisoned Agent Memory via Causal Attribution and Structural Anomaly Detection

arXiv cs.AI · 2026-05-25 Cached

MemAudit is a post-hoc auditing framework for memory-augmented LLM agents that identifies poisoned memories by combining counterfactual influence scores and structural anomaly detection, reducing attack success rates from over 70% to 0% in realistic scenarios.

0 favorites 0 likes
← Back to home

Submit Feedback