Tag
This paper introduces MemoryDecoder at Scale, scaling parametric long-term memory models to 6.9B parameters pretrained on 300B tokens, showing that independently scaling memory is more parameter-efficient than scaling base models alone.