Tag
This paper introduces the MUSE task to evaluate context updating in LLMs and proposes PLUME, a training-free method that improves performance in sequential evolution settings with significant gains on the MUSE-Bench benchmark.