Tag
VoxMem is a benchmark for evaluating multimodal memory in Large Audio Language Models, focusing on acoustic evidence types and multi-session memory operations. It reveals significant gaps in current models, such as poor retention of speaker identity and paralinguistic cues compared to semantic content.
This paper investigates whether LLMs provide grounded pronunciation feedback to L2 English learners, finding that LLMs often rely on stereotypes and prior knowledge rather than acoustic evidence, leading to inaccurate but coherent feedback.