VoxMem: Benchmarking Multimodal Memory in Large Audio Language Models
Summary
VoxMem is a benchmark for evaluating multimodal memory in Large Audio Language Models, focusing on acoustic evidence types and multi-session memory operations. It reveals significant gaps in current models, such as poor retention of speaker identity and paralinguistic cues compared to semantic content.
View Cached Full Text
Cached at: 09/30/26, 08:23 AM
Paper page - VoxMem: Benchmarking Multimodal Memory in Large Audio Language Models
Source: https://huggingface.co/papers/2609.32607
Abstract
Spokenconversationalsystemsmustrecoverinformationfrompriorinteractions(i.e.,memory),yetrelevantinformationinspeechextendsbeyondwhatwassaidtowhosaidit,howitwasspoken,andwhatwasaudible,informationthatexistsonlyintheaudiosignalandcannotberecoveredfromatranscript.Beyondwhattoremember,memoryalsodemandsdiverseoperations:retrievingasinglefact,integratingevidenceacrossturns,trackinganevolvingstate.Realinteractionsfurtherunfoldacrosssessions,meaninginformationaccumulatesacrossdistinctepisodesratherthanasinglecontinuousrecording.Existingbenchmarksfallshortonallthreedimensions:theyfocusprimarilyonlexicalcontent,adoptlimitedandadhocmemoryoperations,andtreatmemoryasasingle-sessionproblem.Wearguethatprincipledmemoryevaluationrequiresjointlycharacterizingtheacousticevidencetoberetainedandtheoperationsappliedtoit,andintroduceataxonomyalongthesetwoaxes.Buildingonthistaxonomy,wepresentVoxMem:3,196evaluationinstancesover34,743spokensessions(177hours)crossingfouracousticevidencetypes(speechsemantics,speakeridentity,paralinguisticcues,environmentalsound)withfourmemoryoperations(informationextraction,multi-sessionreasoning,temporaltracking,andanswerrefusal),groundedinmulti-sessionhistoriesandstratifiedacrosscontextbudgetsfrom8Kto64Ktokens.Evaluating15LALMs,nomodelexceeds40%at32K.Modelsretainwhatwassaidfarbetterthanwhosaidit,how,orwhatwasaudible,agapthatwidensforcomplexoperations,growswithhistorylength,andmanifestsasqualitativelydistinctfailuremodesacrossevidencetypes.VoxMemaimstoprovideafoundationtomeasureanddriveprogressonthefullscopeofspokenconversationalmemory.
View arXiv pageView PDFProject pageGitHub0Add to collection
Get this paper in your agent:
hf papers read 2609\.32607
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2609.32607 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2609.32607 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2609.32607 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
MemLens: Benchmarking Multimodal Long-Term Memory in Large Vision-Language Models
MemLens is a new benchmark for evaluating memory capabilities in large vision-language models through multi-session conversations. It compares long-context and memory-augmented approaches, revealing limitations in both and motivating hybrid architectures.
Beyond Dyadic Memory: Interaction-Aware Multimodal Memory with Adaptive Agentic Retrieval for Multi-Party Spoken Conversations
Proposes VoxPolyMem, an interaction-aware multimodal memory framework with adaptive agentic retrieval for multi-party spoken conversations, achieving state-of-the-art performance on new benchmarks.
EM^2Mem: Event-Centric Multimodal Memory for Large Language Models
EM^2Mem proposes an event-centric multimodal memory framework that binds heterogeneous evidence to event anchors for compact, generation-ready memory in long-video question answering, improving accuracy and reducing latency.
VoiceLongMemEval: Do Assistants Remember How You Sounded?
The paper introduces VoiceLongMemEval (VLME), a benchmark that evaluates AI assistants' ability to remember and reason over paralinguistic metadata like emotion and prosody from voice in long-term conversations, revealing an 'affect gap' in current models.
ChildVox: A Speech, Audio, and Large Audio-Language Model Benchmark in Understanding and Characterizing Sound across Childhood
ChildVox presents a comprehensive benchmark for analyzing children's acoustic communication across developmental stages, integrating over 20 sub-tasks from 17 child-centered audio and speech datasets.