Tag
The paper introduces MemUse, a benchmark that shows direct QA accuracy for conversational memory does not predict user satisfaction, whereas natural integration of prior context does, revealing a large gap between recall and conversational use.