Tag
A paper proposing a mechanism for enterprise agents to improve by safely converting messy daily work into learning data, using a data proxy and control layer, with AREAL2.0 demonstrating online RL from real interaction traces.
Introduces Supersede, an environment to diagnose and train the memory-update gap in LLM agents, showing that standard models fail to maintain current facts as conversations grow, and that GRPO fine-tuning can improve performance.
Proposes Erase-then-Delta Attention (EDA), a memory update rule for linear attention that decouples erase and write addresses to selectively suppress stale information before writing new content. Experiments on 2.5B dense and 25B MoE models demonstrate consistent gains in standard and long-context evaluations.
ChatGPT's memory update led to the chatbot fixating on users' painful personal details, causing mental health crises and at least 20 lawsuits against OpenAI, including one after a user's suicide.