Tag
This paper introduces a paired-history audit method to evaluate update sufficiency in context compression for AI memory systems, revealing failures where memories answer correctly now but cannot handle future updates.
This paper demonstrates that scaling point-in-time language models—trained exclusively on text available up to each calendar date—can substantially narrow the performance gap with unrestricted models, enabling valid backtests and causal inference in finance and social sciences. The authors train decoder-only transformers up to 4B parameters on 1 trillion chronologically filtered tokens and release the full pipeline.
This paper introduces MemStrata, a retrieval memory system that maintains temporal validity to eliminate stale-fact errors in AI agents over evolving knowledge. It outperforms RAG on evolving benchmarks while preserving static recall, using a deterministic supersession layer without LLM calls.