Tag
ArborMem introduces an online memory framework for large language models that represents conversations as a navigable forest of interaction states, outperforming baselines on memory benchmarks and introducing BranchMemEval as a new diagnostic benchmark.
The study compares MemoryLake, a structured memory backend, with other systems on the MemoryArena benchmark, finding it has higher success rates in some tasks but not establishing a clear advantage across all domains.
The project introduces long-term personality and memory persistency for AI companion agents, featuring mechanisms like evolving personalities, grudge-holding, and associative memory web.
Presents EvoGraph-Mem, a failure-aware editable graph memory framework for long-term language agents that tracks positive/negative evidence and activation states for insights, enabling memory maintenance through utility-aware retrieval and graph-level editing.
LycheeMemory V2 introduces semantic segment-level consolidation for LLM agent memory, reducing construction costs while improving or maintaining accuracy on LoCoMo and LongMemEval benchmarks.
Introduces Controlled Memory Interference (CMI), a diagnostic framework for studying how LLM agent memory evolves under different memory relationships, revealing that relationship-specific interference suppresses update plasticity and that interference-aware training improves valid update distinction.
MobileMem introduces a benchmark and framework for on-device AI systems to learn from year-long mobile experiences, focusing on temporal reasoning, knowledge updating, and preference inference.
After leaving Shanda Innovation Institute, Lin Kaisen founded a new company focused on long-term memory agents, and mentioned his experience working under the guidance of Boss Chen.
AgentMemBench is a systematic benchmark that evaluates five long-term memory management strategies for conversational AI agents across three datasets, finding that external key-value store retrieval dominates on quality but incurs a larger memory footprint.
An essay arguing that compression, not longer context windows, is the load-bearing primitive for long-term AI memory and relational continuity, drawing an analogy to compute and storage in computing.
This paper introduces IFCMemoryBench, a human-validated benchmark for evaluating long-term memory in LLM-based agents for BIM information retrieval. It shows that current memory systems achieve only 32.4% answer accuracy, revealing a domain-transfer gap in agent memory.
The HKU team open-sourced DeepTutor, a personalized AI learning companion with long-term memory, multiple knowledge base retrieval methods, and a skill ecosystem, capable of tutoring, problem solving, question generation, research, visualization, and learning path planning.
This paper introduces MemoryDecoder at Scale, scaling parametric long-term memory models to 6.9B parameters pretrained on 300B tokens, showing that independently scaling memory is more parameter-efficient than scaling base models alone.
Explores different long-term memory architectures for AI agents, with a focus on an agent-as-memory-controller approach using neon postgres.
MOSAIC is a structured, conflict-aware long-term memory framework for LLM agents that uses entity-typed graph storage, hash-accelerated retrieval, and active conflict detection to achieve high accuracy and efficiency on long-conversation QA and factual conflict detection tasks.
Introduces MemOps, a benchmark that reformulates conversational memory as lifecycle operations with structured traces, enabling operation-level diagnosis of memory failures in LLM-based agents, revealing that current systems are far from uniformly reliable.
ReflectWorld-MM is an entity-oriented multimodal memory system for open-ended video streams, using hierarchical long-term memory inspired by human memory theory, achieving state-of-the-art accuracy on six benchmarks.
Amazon has announced an agentic version of Alexa with long-term memory and integration with over 1,000 apps, enhancing its capabilities as a personal AI assistant.
NapMem is a framework that treats long-term user memory as a structured action space rather than passive retrieval, using a multi-granularity memory pyramid and reinforcement learning to train agents to navigate memory. Experiments show competitive performance on memory-intensive tasks.
A reproducible demo showing how adding long-term memory to CrewAI with Mem0 reduces token usage and latency via a fast/deep routing heuristic, with real measurements and a live dashboard.