Kept seeing complaints about memory benchmarks, so I built one. Glasshouse v0.1 is out
Summary
Glasshouse v0.1 is a memory benchmark tool designed to address issues with existing AI memory benchmarks by providing multi-language support, varied conversation lengths, and detailed evaluation across multiple axes.
Similar Articles
My memory benchmark is close to done, looking for companies and devs to help build it out from here
A near-complete memory benchmark tool named Glasshouse is being developed, with public rules and structure, seeking community contributions to build out the benchmark files and ensure fair evaluation.
I built a benchmark for AI “memory” in coding agents. looking for others to beat it.
Developer created a new benchmark called continuity-benchmarks to test AI coding agents' ability to maintain consistency with project rules during active development, addressing gaps in existing memory benchmarks that focus on semantic recall rather than real-time architectural consistency and multi-session behavior.
@witcheer: someone in the community worked on a real benchmark of seven self-hostable memory providers for Hermes Agent, each fed …
A community member benchmarked seven self-hostable memory providers for Hermes Agent, testing each with 71,060 conversation turns and 3,750 questions about changing facts; full results and GitHub repo are in the thread.
AgentMemBench: A Systematic Benchmark for Evaluating Long-Term Memory Management Strategies in Conversational AI Agents
AgentMemBench is a systematic benchmark that evaluates five long-term memory management strategies for conversational AI agents across three datasets, finding that external key-value store retrieval dominates on quality but incurs a larger memory footprint.
MemoBench: Benchmarking World Modeling in Dynamically Changing Environments
MemoBench is a diagnostic benchmark for evaluating video generation models' memory consistency in dynamically changing environments, where objects disappear and reappear in updated states. It includes 360 ground-truth clips and an evaluation suite combining automated metrics with VQA-based assessment, revealing insights into memory consistency challenges.