long-horizon-evaluation

Tag

Cards List
#long-horizon-evaluation

MemEvoBench: Benchmarking Memory MisEvolution in LLM Agents

arXiv cs.CL · 2026-04-20 Cached

MemEvoBench introduces the first benchmark for evaluating memory safety in LLM agents, measuring behavioral degradation from adversarial memory injection, noisy outputs, and biased feedback across QA and workflow tasks. The work reveals that memory evolution significantly contributes to safety failures and that static defenses are insufficient.

0 favorites 0 likes
← Back to home

Submit Feedback