VoxMem: Benchmarking Multimodal Memory in Large Audio Language Models

Hugging Face Daily Papers Papers

Summary

VoxMem is a benchmark for evaluating multimodal memory in Large Audio Language Models, focusing on acoustic evidence types and multi-session memory operations. It reveals significant gaps in current models, such as poor retention of speaker identity and paralinguistic cues compared to semantic content.

Spoken conversational systems must recover information from prior interactions (i.e., memory), yet relevant information in speech extends beyond what was said to who said it, how it was spoken, and what was audible, information that exists only in the audio signal and cannot be recovered from a transcript. Beyond what to remember, memory also demands diverse operations: retrieving a single fact, integrating evidence across turns, tracking an evolving state. Real interactions further unfold across sessions, meaning information accumulates across distinct episodes rather than a single continuous recording. Existing benchmarks fall short on all three dimensions: they focus primarily on lexical content, adopt limited and ad hoc memory operations, and treat memory as a single-session problem. We argue that principled memory evaluation requires jointly characterizing the acoustic evidence to be retained and the operations applied to it, and introduce a taxonomy along these two axes. Building on this taxonomy, we present VoxMem: 3,196 evaluation instances over 34,743 spoken sessions (177 hours) crossing four acoustic evidence types (speech semantics, speaker identity, paralinguistic cues, environmental sound) with four memory operations (information extraction, multi-session reasoning, temporal tracking, and answer refusal), grounded in multi-session histories and stratified across context budgets from 8K to 64K tokens. Evaluating 15 LALMs, no model exceeds 40% at 32K. Models retain what was said far better than who said it, how, or what was audible, a gap that widens for complex operations, grows with history length, and manifests as qualitatively distinct failure modes across evidence types. VoxMem aims to provide a foundation to measure and drive progress on the full scope of spoken conversational memory.
Original Article
View Cached Full Text

Cached at: 09/30/26, 08:23 AM

Paper page - VoxMem: Benchmarking Multimodal Memory in Large Audio Language Models

Source: https://huggingface.co/papers/2609.32607

Abstract

Spokenconversationalsystemsmustrecoverinformationfrompriorinteractions(i.e.,memory),yetrelevantinformationinspeechextendsbeyondwhatwassaidtowhosaidit,howitwasspoken,andwhatwasaudible,informationthatexistsonlyintheaudiosignalandcannotberecoveredfromatranscript.Beyondwhattoremember,memoryalsodemandsdiverseoperations:retrievingasinglefact,integratingevidenceacrossturns,trackinganevolvingstate.Realinteractionsfurtherunfoldacrosssessions,meaninginformationaccumulatesacrossdistinctepisodesratherthanasinglecontinuousrecording.Existingbenchmarksfallshortonallthreedimensions:theyfocusprimarilyonlexicalcontent,adoptlimitedandadhocmemoryoperations,andtreatmemoryasasingle-sessionproblem.Wearguethatprincipledmemoryevaluationrequiresjointlycharacterizingtheacousticevidencetoberetainedandtheoperationsappliedtoit,andintroduceataxonomyalongthesetwoaxes.Buildingonthistaxonomy,wepresentVoxMem:3,196evaluationinstancesover34,743spokensessions(177hours)crossingfouracousticevidencetypes(speechsemantics,speakeridentity,paralinguisticcues,environmentalsound)withfourmemoryoperations(informationextraction,multi-sessionreasoning,temporaltracking,andanswerrefusal),groundedinmulti-sessionhistoriesandstratifiedacrosscontextbudgetsfrom8Kto64Ktokens.Evaluating15LALMs,nomodelexceeds40%at32K.Modelsretainwhatwassaidfarbetterthanwhosaidit,how,orwhatwasaudible,agapthatwidensforcomplexoperations,growswithhistorylength,andmanifestsasqualitativelydistinctfailuremodesacrossevidencetypes.VoxMemaimstoprovideafoundationtomeasureanddriveprogressonthefullscopeofspokenconversationalmemory.

View arXiv pageView PDFProject pageGitHub0Add to collection

Get this paper in your agent:

hf papers read 2609\.32607

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2609.32607 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2609.32607 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2609.32607 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

VoiceLongMemEval: Do Assistants Remember How You Sounded?

arXiv cs.AI

The paper introduces VoiceLongMemEval (VLME), a benchmark that evaluates AI assistants' ability to remember and reason over paralinguistic metadata like emotion and prosody from voice in long-term conversations, revealing an 'affect gap' in current models.