Tag
Garry Tan announced that GBrain now supports Jev, improving its remembering/dream cycle while keeping retrieval behavior unchanged.
Mnemon is a memory agent that keeps conversations as raw dated records and divides memory work into System 1 fast judgments by a small decision model (Jev) and System 2 planning by an LLM, achieving state-of-the-art LoCoMo (91.7%) and LongMemEval-S (94.4%) scores at low token cost.
MemLife introduces a multimodal memory system for long-term egocentric videos that builds entity-grounded text episodes and retrieves them with a time-indexed agentic reader, improving over training-free baselines by 4.6-12.0% on long-horizon benchmarks. A reinforcement learning framework called MemOpt further optimizes the memory writer for faithfulness and retrievability, yielding consistent 2.7-5.0% gains.
The blog post introduces a multi-part series on techniques for efficient continuous learning in AI, exploring alternatives to reinforcement learning and using KV cache memory at Ramp Labs.
This paper introduces Memory of Memory (MoM), a framework for LLM agent memory that commits current values on arrival while retaining displaced values as provenance, improving accuracy and reducing stale answers.
This paper measures how AI agents contaminate stores they write to, identifying a threshold for error propagation and validating findings with synthetic and real-world Wikidata data.
This article explains the concept of world models in AI agents, highlighting their role in tracking evolving states and transitions, distinct from context or memory, with examples like narrative world models for long-form fiction.
El artículo explora la investigación en Proyecto Chatty sobre la construcción de agentes de IA con continuidad, enfocándose en el sistema alrededor del modelo, incluyendo contexto, memoria, herramientas y permisos.
This paper proposes a hierarchical architecture for long-horizon AI agents, incorporating levels, ticks, and cascaded intelligence to enable continual operation without forgetting, demonstrated over a ten-day campaign.
The author critiques current AI agent memory systems for following old data engineering patterns, highlighting the risk of gaps in memory and advocating for auditable raw trajectories before layered summaries.
ThinkFlow is a novel end-to-end latent memory framework for lifelong conversational agents that uses probabilistic vectors to overcome textual memory bottlenecks. It enables autonomous personalization through self-evolution and test-time learning, outperforming existing memory systems.
EchoPath introduces a model-agnostic memory system for GUI agents that replays validated execution trajectories, significantly reducing token cost and execution time for enterprise recurrent tasks.
The article explains why knowledge graphs are superior to flat lists for memory storage in AI systems, as they effectively handle multiple references to the same entity and prevent fragmented or contradictory search results.
The author tested a memory startup's product against their OpenClaw agent using a simple MEMORY.md setup, finding that the startup's runtime couldn't handle temporal data while their markdown-based approach worked. They open-sourced the setup and shared guidelines for agent memory management.
SimSkill is a lifelong learning AI agent that autonomously masters traffic simulation by identifying capability gaps, generating tasks, and using memory systems to improve performance, showing up to 25% improvement in task completion on benchmarks.
The post questions whether memory in AI agents is essential or a workaround for poor architecture, and asks practitioners what needs to be persisted in production systems.
The article explores designing memory systems for AI agents that are auditable by humans, suggesting fields like provenance, scope, and expiration rules to maintain clarity and prevent stale information.
The AQuA v2 preprint introduces a memory system for research agents that uses persistent evidence from accepted and rejected experiments to guide future proposals, emphasizing the importance of evidence lifecycle management.
The author questions whether existing agent-memory setups track the success of learned workflows and proposes their approach of using counters and evolution logs, seeking prior art or conventions in the field.
Multiple AI Agent tools and frameworks on GitHub's trending list, such as mattpocock/skills and obra/superpowers, have garnered widespread attention due to their innovation and practicality, marking a shift in AI Agent development from talking agents to infrastructure with skill standards, persistent context, and multi-agent collaboration.