Memory Decoder at Scale: A Pretrained, Parametric Long-Term Memory

Hugging Face Daily Papers Papers

Summary

This paper introduces MemoryDecoder at Scale, scaling parametric long-term memory models to 6.9B parameters pretrained on 300B tokens, showing that independently scaling memory is more parameter-efficient than scaling base models alone.

Decoder-only language models entangle long-term memory and reasoning in a single parameter set, making it difficult to scale memory capacity independently. Memory Decoder introduces a parametric long-term memory module but only studies it at a relatively small scale. In this work, we present Memory Decoder at Scale, scaling memory models up to 6.9B parameters and pretraining them on 300B tokens. At this data scale, the combined cost of indexing and search makes a standard Faiss pipeline infeasible. We address this bottleneck with a distributed pipeline for Faiss indexing and retrieval, together with sparse, batch-wise loading of kNN distributions. Across model scales, we find that allocating more parameters to memory yields a better parameter-performance tradeoff than scaling the base model alone. On 17 benchmarks, pairing a 6.9B general memory with Pythia-410M raises its average score from 29.86 to 37.34, surpassing Pythia-12B (37.24) with 39% fewer total parameters. For Qwen3 Base models ranging from 0.6B to 14B, 1.7B domain memories improve the average score across the three domains by more than 9 points at every scale. Overall, our results demonstrate that independently scaling pretrained memory offers a more parameter efficient path to improving language model performance.
Original Article
View Cached Full Text

Cached at: 07/31/26, 05:52 AM

Paper page - Memory Decoder at Scale: A Pretrained, Parametric Long-Term Memory

Source: https://huggingface.co/papers/2607.27919

Abstract

Decoder-onlylanguagemodelsentanglelong-termmemoryandreasoninginasingleparameterset,makingitdifficulttoscalememorycapacityindependently.MemoryDecoderintroducesaparametriclong-termmemorymodulebutonlystudiesitatarelativelysmallscale.Inthiswork,wepresentMemoryDecoderatScale,scalingmemorymodelsupto6.9Bparametersandpretrainingthemon300Btokens.Atthisdatascale,thecombinedcostofindexingandsearchmakesastandardFaisspipelineinfeasible.WeaddressthisbottleneckwithadistributedpipelineforFaissindexingandretrieval,togetherwithsparse,batch-wiseloadingofkNNdistributions.Acrossmodelscales,wefindthatallocatingmoreparameterstomemoryyieldsabetterparameter-performancetradeoffthanscalingthebasemodelalone.On17benchmarks,pairinga6.9BgeneralmemorywithPythia-410Mraisesitsaveragescorefrom29.86to37.34,surpassingPythia-12B(37.24)with39%fewertotalparameters.ForQwen3Basemodelsrangingfrom0.6Bto14B,1.7Bdomainmemoriesimprovetheaveragescoreacrossthethreedomainsbymorethan9pointsateveryscale.Overall,ourresultsdemonstratethatindependentlyscalingpretrainedmemoryoffersamoreparameterefficientpathtoimprovinglanguagemodelperformance.

View arXiv pageView PDFProject pageGitHub0Add to collection

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2607.27919 in a model README.md to link it from this page.

Datasets citing this paper1

#### Rubin-Wei/MemoryDecoder-at-Scale-domain-data Viewer• Updatedabout 4 hours ago • 11M • 9

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2607.27919 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

Hidden Decoding at Scale: Latent Computation Scaling for Large Language Models

arXiv cs.CL

This paper introduces Hidden Decoding, a sequence-length scaling method for LLMs that adds internal computation per token by expanding each token into multiple streams with independent embeddings, using Stream-Factorized Attention to keep costs low. Experiments on models up to 617B parameters show consistent improvements over baselines, demonstrating a practical fixed-backbone scaling path.

δ-mem: Efficient Online Memory for Large Language Models

Hugging Face Daily Papers

The paper introduces δ-mem, a lightweight memory mechanism that enhances large language models by augmenting a frozen attention backbone with a compact associative memory state. It demonstrates improved performance on memory-heavy benchmarks with minimal computational overhead.

Understanding Is Done Early: A Depth Division of Labor in Large Language Models and Its Use for Unbounded-Context Memory

arXiv cs.CL

This paper introduces CoMem, a method that exploits the depth-wise division of labor in LLMs to cache intermediate residual tensors and recompute only upper layers for retrieval, enabling bounded read compute and memory independent of stored-context length. Evaluated on Qwen3-8B, CoMem achieves strong long-context performance with significant memory savings and prefill speedups.

ConvMem: Convolutional Memory for Long-Context Reasoning

arXiv cs.AI

ConvMem is a training-free, parallelizable framework that reformulates long-context reasoning in large language models as hierarchical convolution to improve efficiency, avoid overfitting, and outperform baseline methods.