Tag
Introduces MAP-Graph, a provenance-aware shared memory layer for multi-agent workflows that filters by permissions, reranks by path trust, and gates risky actions, achieving 94.96% task success in controlled benchmarks.
A best-paper research from NVIDIA and collaborators introduces WorldTrace, a training-free framework that keeps compressed memory addressable in autoregressive video world models by assigning fixed slot-rank positions, enabling coherent long rollouts and long-range recall beyond the training horizon.
A photographer reflects on finding a job at Polo Ralph Lauren through a newspaper classified ad in 2000, using digital archives to verify his memory.
A new preprint called TEPA treats memory validity as a first-class state, revoking outdated precedents when new evidence conflicts while keeping audit trails. It outperforms append-only and last-write-wins in a complete-reversal experiment, though results are not yet independently reproduced.
This arXiv survey (1,547 papers, 2024-2026) systematically maps the field of long-horizon LLM agents, disambiguating long-horizon, long-context, and long-term memory, and organizing research into six lifecycle categories while identifying the core 'horizon gap' and open measurement problems.
A paper explores 'memory provenance laundering' in LLM agents, where long-term memory can turn untrusted observations into seemingly trusted context, and proposes preserving provenance through memory consolidation.
This paper introduces ReMEMBER, a missing-evidence memory framework for streaming dialogue summarization that retrieves and refines evidence from long histories to resolve gaps in current windows under fixed memory budgets, along with a benchmark for evaluation.
The paper 'Memory Reward Inflation in Self-Improving LLM Agents' shows that self-improving agents with frozen weights can still degrade by trusting flawed LLM-generated memory scores, with models endorsing 31–54% of their own wrong answers. This 'Echo Gap' persists across stronger LLMs.
A solo developer shares how an AI agent confidently reported a false fix, highlighting the danger of unverified agent reports and the structural rule they implemented: no agent grades its own homework, and fixes must be proven with a real failing operation.
A Reddit user compares the cheapest hardware options for achieving 128GB+ memory for local AI in 2026, covering used GPUs, unified memory systems, and cloud alternatives.
Neo4j's Will Lyon presents a recorded session on designing stateful AI agents with memory and context graphs, covering pitfalls and integrations with Salesforce Agentforce and Databricks.
This paper introduces MERIT, a training-free agent that uses causal episodic memory of past repair outcomes to improve subsequent Text-to-SQL generations, boosting execution accuracy on Spider and BIRD benchmarks.
ContextWeave is a new longitudinal benchmark that evaluates whether recalled memory improves downstream agent performance in realistic office-work streams, using privacy-preserved multi-month workflows of 14 participants to create 1,005 executable tasks.
Mem Agent now lets users see its internal tracking of projects, tasks, reminders, and routines, offering transparency to help align the AI assistant's mental model with the user's world. Available in the Mem Proactive plan.
Hansel is a product that helps you remember everything you've worked on, likely a memory or productivity tool.
PI-Mem is a parallel-iterative memory mechanism that pushes long-context reasoning to 3.6M tokens, outperforming recurrent-memory baselines while achieving significant inference speedups.
MemArena is a new ego-centric benchmark for evaluating on-device personal memory assistants, using a MASim agent simulator to generate multi-session conversational worlds and ground truth across recall, reasoning, and trustworthiness dimensions. Initial results show memory-backend choice often matters more than reader scale, and permission-aware access remains a universal challenge.
SK Hynix introduces CMM-Ax, an ASIC-based CXL-PNM solution developed with Marvell Technology, designed to overcome memory bottlenecks in long-context LLM inference, achieving up to 5.5× higher throughput than GPU-only systems.
The author argues that AI agent recall is largely solved but permission/authorization over retrieved memories is the missing layer, introducing Provem, an open-source governance layer between agents and memory stores with benchmarks showing it eliminates compliance violations.
This paper introduces MemoryForge, a framework for synthesizing lifelong autobiographical memory from brief target personas to enable frozen LLMs to exhibit more human-like behaviors in role-play and user-simulation, outperforming descriptive conditioning baselines.