semantic-compression

Tag

Cards List
#semantic-compression

SeKV: Resolution-Adaptive KV Cache with Hierarchical Semantic Memory for Long-Context LLM Inference

arXiv cs.CL · 2026-07-01 Cached

SeKV is a resolution-adaptive KV cache method that organizes context into entropy-guided semantic spans stored across a GPU-CPU hierarchy, enabling selective token-level reconstruction during decoding while reducing GPU memory by 53.3% versus full caching at 128K context.

0 favorites 0 likes
#semantic-compression

What if context compression is a diffusion noise function? Proposal + honest results from untrained-model experiments [R]

Reddit r/MachineLearning · 2026-06-26

Proposes treating semantic compression as a diffusion noise function for handling massive context beyond model windows, using multi-pass reading at decreasing compression levels. Untrained-model experiments show components work in isolation but the full chain needs training to resolve binding bottleneck.

0 favorites 0 likes
#semantic-compression

SimpleMem: Efficient Lifelong Memory for LLM Agents

Papers with Code Trending · 2026-01-05 Cached

Introduces SimpleMem, an efficient memory framework for LLM agents that uses semantic lossless compression to improve accuracy and reduce token consumption, achieving 26.4% F1 improvement and up to 30x reduction in inference-time token usage.

0 favorites 0 likes
← Back to home

Submit Feedback