multimodal-memory

Tag

Cards List
#multimodal-memory

EpiCon: Collective Agent Learning through Co-Evolving Multimodal Memory

Hugging Face Daily Papers ↗ · 2d ago Cached

EpiCon presents a shared multimodal memory framework for collective learning among AI agents, enhancing performance across eleven benchmarks without updating host model parameters.

0 favorites 0 likes
#multimodal-memory

VoxMem: Benchmarking Multimodal Memory in Large Audio Language Models

Hugging Face Daily Papers ↗ · 5d ago Cached

VoxMem is a benchmark for evaluating multimodal memory in Large Audio Language Models, focusing on acoustic evidence types and multi-session memory operations. It reveals significant gaps in current models, such as poor retention of speaker identity and paralinguistic cues compared to semantic content.

0 favorites 0 likes
#multimodal-memory

Beyond Dyadic Memory: Interaction-Aware Multimodal Memory with Adaptive Agentic Retrieval for Multi-Party Spoken Conversations

Hugging Face Daily Papers ↗ · 5d ago Cached

Proposes VoxPolyMem, an interaction-aware multimodal memory framework with adaptive agentic retrieval for multi-party spoken conversations, achieving state-of-the-art performance on new benchmarks.

0 favorites 0 likes
#multimodal-memory

Parametric Multimodal User Memory: Storing What Captions Cannot Carry

arXiv cs.CL ↗ · 2026-09-01 Cached

This research paper introduces a parametric multimodal user memory system that improves AI agents' ability to recall users by integrating perceptual data like voice and appearance, using vision-language models and dedicated encoders to surpass text-based methods.

0 favorites 0 likes
#multimodal-memory

EM^2Mem: Event-Centric Multimodal Memory for Large Language Models

Hugging Face Daily Papers ↗ · 2026-09-01 Cached

EM^2Mem proposes an event-centric multimodal memory framework that binds heterogeneous evidence to event anchors for compact, generation-ready memory in long-video question answering, improving accuracy and reducing latency.

0 favorites 0 likes
#multimodal-memory

Beyond Retrieval: Analytic Memory for Multimodal Agents

arXiv cs.AI ↗ · 2026-08-03 Cached

This paper introduces AdaMM, a framework that complements retrieval-based multimodal memory with analytic memory, enabling filtering, aggregation, ranking, and temporal comparison over accumulated observations. Experiments on MemEye and MemGallery benchmarks show improvements of up to 11.3% and 7.3% respectively.

0 favorites 0 likes
#multimodal-memory

ViSAGE: Constructing Self-Correcting Memories for Long-Form Video Understanding

arXiv cs.AI ↗ · 2026-08-03 Cached

ViSAGE is a multimodal agentic memory framework for long-form video understanding that builds self-correcting, entity-centric memories via cross-modal binding, bidirectional memory refinement, and multi-agent cross-verification, achieving 5.9% higher accuracy than baselines.

0 favorites 0 likes
← Back to home

Submit Feedback