@yoheinakajima: ran my first benchmark this weekend (longmemeval) mostly to test activegraph, learned a lot! - this is a stepping stone…

X AI KOLs Timeline Models

Summary

Yohei Nakajima ran the LongMemEval benchmark on ActiveGraph, achieving 85.6% QA accuracy and 86.2% turn answer-in-context, demonstrating the effectiveness of event-based agent systems for long-term memory.

ran my first benchmark this weekend (longmemeval) mostly to test activegraph, learned a lot! - this is a stepping stone to show the event based agent system works. the AI convinced me not to start with graph extraction of facts/entities - learned running benchmarks takes a long time - understand more what a good benchmark vs bad benchmark is (I think this is pretty thorough) - fully reproducible open source repo and tests - seems like ActiveGraph is well suited for this, performed solid, which is a good start
Original Article
View Cached Full Text

Cached at: 05/26/26, 07:05 AM

ran my first benchmark this weekend (longmemeval) mostly to test activegraph, learned a lot!

  • this is a stepping stone to show the event based agent system works. the AI convinced me not to start with graph extraction of facts/entities
  • learned running benchmarks takes a long time
  • understand more what a good benchmark vs bad benchmark is (I think this is pretty thorough)
  • fully reproducible open source repo and tests
  • seems like ActiveGraph is well suited for this, performed solid, which is a good start

no extraction here, just deterministic ingestion as a first test

for now just research, but i’m also testing cofounder from @intelligenceco which is handling the website/newsletter/blog which has been a fun way to add some polish

Awesome

that’s so cool, please do! i can really only work on this on weekends so would love to see other ppl push it faster

Similar Articles