longmemeval

Tag

Cards List
#longmemeval

What would a memory benchmark have to do before you'd trust the number?

Reddit r/AI_Agents · 2026-08-12

The author critiques the reliability of memory benchmarks for LLMs, citing issues like incorrect golden answers in LoCoMo and inconsistent scoring metrics yielding large score gaps. They question whether practitioners trust these numbers and discuss recent improvements like LongMemEval's knowledge update questions.

0 favorites 0 likes
#longmemeval

@yoheinakajima: in arxiv paper #2, i tackle the last topic from paper #1: @activegraphai as an architectural affordance for self-improv…

X AI KOLs Following · 2026-06-10 Cached

This paper introduces Regimes, an auditable, held-out-gated improvement loop built on the ActiveGraph runtime for self-improving agents. It demonstrates modest improvements on the LongMemEval dataset by autonomously discovering prompt repairs that pass static checks, sandbox execution, and held-out validation.

0 favorites 0 likes
#longmemeval

Regimes: An Auditable, Held-Out-Gated Improvement Loop Demonstrated on LongMemEval with ActiveGraph

arXiv cs.AI · 2026-06-10 Cached

Regimes is an auditable, held-out-gated improvement loop built on the ActiveGraph event-sourced runtime. It diagnoses failures in AI agents, proposes repairs, and promotes them only after passing multiple gates, improving accuracy on LongMemEval by up to +0.10.

0 favorites 0 likes
#longmemeval

@yoheinakajima: i know it's backwards order, but experiment #2:

X AI KOLs Timeline · 2026-06-01 Cached

In our second longmemeval experiment, we introduce semantic ingestion into recall leveraging the ActiveGraph runtime, improving retrieval from 60.6% to 83.4%/84.8% for flat/agentic retrieval with LLM ingestion.

0 favorites 0 likes
#longmemeval

#1 on memory benchmark LongMemEval with Gemini Flash, not Pro [R]

Reddit r/MachineLearning · 2026-05-17

A novel memory retrieval system inspired by episodic memory theory achieves state-of-the-art 96.4% top-50 accuracy on the LongMemEval benchmark using Gemini Flash, outperforming larger Pro-based baselines by isolating retrieval quality from model capability.

0 favorites 0 likes
← Back to home

Submit Feedback