Task-Focused Memorization for Multimodal Agents
Summary
Introduces TaskMem, a reinforcement-learning-based framework for dynamic memorization in multimodal agents, achieving accuracy improvements of 6.3%, 7.0%, and 5.3% on streaming video benchmarks.
View Cached Full Text
Cached at: 06/01/26, 03:17 AM
Paper page - Task-Focused Memorization for Multimodal Agents
Source: https://huggingface.co/papers/2605.31075
Abstract
A reinforcement-learning-based framework called TaskMem is introduced to dynamically determine what information to store in long-term memory for multimodal agents, improving performance on streaming video benchmarks.
Long-term memoryis essential formultimodal agentsto build coherent experience, accumulate world knowledge, and achievecontinual learning. However, constructing effective memory goes beyond memory module design and basic requirements such as accuracy and fidelity; the key challenge lies in determining what to memorize.Multimodal agents, such as embodied agents, continuously perceive, reason, and act in real or virtual environments, receiving an unbounded stream of multimodal observations. From this combinatorial explosion of information, an agent must selectively retain content that is relevant to its role in the environment and valuable for future tasks. To bridge this gap, we frame memory generation as a learnablememorization policyand introduce TaskMem (Task-focused MemorizationPolicy Learning), a reinforcement-learning-based framework that enables the policy to dynamically adjust its focus to the demands of real tasks encountered in the environment. TaskMem adopts atwo-phase trainingparadigm: Phase One learns how to memorize by optimizingmemory qualityunder fundamental fidelity requirements; Phase Two occurs after deployment, where the agent learns what to memorize by tuning an adapter on its base MLLM, using recent environment tasks to define areward modelthat guides thememorization policytoward task-relevant content. To evaluate our approach, we reformulateVideoMME,EgoLife, andEgoTempointostreaming benchmarksthat simulate a realistic setting in which an agent processes streaming observations and handles tasks arriving online. To isolate memory assessment, the questions must be answered using only the agent’s memory, without access to raw video. Built onQwen3-VL-30B-A3B, TaskMem improvesVQA accuracyby 6.3%, 7.0%, and 5.3% on these benchmarks, respectively.
View arXiv pageView PDFProject pageAdd to collection
Get this paper in your agent:
hf papers read 2605\.31075
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2605.31075 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2605.31075 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2605.31075 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
@maximelabonne: Such a cool collab, very happy to see this kind of recipe getting open-sourced Train LFM2.5-2.6B on all the harnesses!
Maximilien Labonne highlights an open-sourced guide to training models with RL inside real agent harnesses like Claude Code, Codex, and OpenCode, enabling training of any model (e.g., LFM2.5-2.6B) on any task set across harnesses.
@omarsar0: Small models are getting really good at reasoning. It's exciting because SLMs can unlock so much at the harness layer. …
TheWebAI released TwIL-LM3-Pro, a 3.6B-parameter open-source SLM that runs locally and achieves 95.4 on BIG-Bench Hard, outperforming Qwen3-8B's 63.7, using a post-training recipe of formal-logic fine-tuning, weight merging, and RL with a programmatic verifier.
@bkdgiffug: Learning Agent is most likely to get stuck at this step: you understand all the concepts, but when it's time to actuall…
LangChain's official "Agents From Scratch" guide offers a hands-on, four-step roadmap for building an email agent with long-term memory and human-machine collaboration, covering LangGraph workflows, LLM-as-Judge testing, human-in-the-loop confirmation, and LangGraph Store memory integration, with notebooks and source code included.
With most information hidden, the game Stratego
A team from CMU, MIT, NYU, and Stanford created Ataraxos, an AI that beat the best Stratego player in history 15-1-4 while training on just 16 GPUs, overcoming the game's massive hidden information and long-horizon challenges that stumped DeepMind's DeepNash.
Clef: our open-source decision models
Cloudflare 发布了两个开源决策模型 Clef 和 Clef-flash,托管于 Workers AI,采用 Apache 2.0 许可并在 Hugging Face 上开放下载,主打低成本、快速且一致的结构化输出,目前在 Jev Decision Index 榜单领先;同时推出新的强化学习微调平台,支持客户针对自身用例对 Clef 进行微调。