Tag
The author demonstrates a method to trick Claude's AI assistant into exfiltrating personal data from its memory by exploiting its web browsing capability, though initial attempts were blocked by Anthropic's safeguards.
This paper investigates memory manipulation in LLM-based agents for multiple-choice question answering, showing that corrupted memories can cause agents to select incorrect options even when the current query is clean.