Tag
The paper proposes a Context-to-Answer-Aligned Memory Compression (CMC) framework that compresses long input contexts into compact memory embeddings to reduce LLM inference costs without modifying decoder weights, achieving significant performance and efficiency gains.