memory-efficiency

Tag

Cards List
#memory-efficiency

PGD-NO: A Neural Operator with Precomputed Geometry Decomposition for 3D Million-scale Physics Simulations

arXiv cs.LG ↗ · 2026-07-10 Cached

PGD-NO is a neural operator that precomputes geometry decomposition to achieve linear memory scalability, enabling high-fidelity physics simulations on meshes exceeding 10 million nodes and overcoming the single-node memory bottleneck.

0 favorites 0 likes
#memory-efficiency

DepthWeave-KV: Token-Adaptive Cross-Layer Residual Factorization for Long-Context KV Cache Compression

arXiv cs.AI ↗ · 2026-07-08 Cached

DepthWeave-KV is a token-adaptive cross-layer residual factorization method for compressing KV cache in long-context transformer inference, achieving 8.3x memory reduction and 72.8 tokens/s at 64K context while preserving near-full-cache task quality across benchmarks.

0 favorites 0 likes
#memory-efficiency

FreqDepthKV: Frequency-Guided Depth Sharing for Robust KV Cache Compression in Long-Context LLM Inference

arXiv cs.AI ↗ · 2026-07-08 Cached

Introduces FreqDepthKV, a frequency-guided depth sharing method for KV cache compression in long-context LLM inference, which factorizes adjacent-layer KV states into shared low-frequency components and sparse high-frequency residuals, improving memory efficiency and throughput while preserving accuracy on benchmarks.

0 favorites 0 likes
#memory-efficiency

@MaxForAI: I agree with Andrew — a major breakthrough in memory efficiency is coming, which is actually a direction that Infra has been working hard on for a long time. And there have already been significant results, such as cache hit. DeepSeek, as early as after the deployment of disk-based KV Cache, cut the input price for cache hits to 1/… of that for cache misses.

X AI KOLs Timeline ↗ · 2026-07-01 Cached

Technical experts discuss the upcoming major breakthrough in memory efficiency, mentioning that DeepSeek has reduced the input price for cache hits to 1/10 to 1/50 of the price for cache misses through KV cache optimization, and reveal that OpenAI engineers have used multiple optimization techniques to cut inference costs by more than half.

0 favorites 0 likes
#memory-efficiency

WAL-RUS: a Rust Rewrite of WAL-G for PostgreSQL Backups

Hacker News Top ↗ · 2026-06-27 Cached

ClickHouse Cloud announces WAL-RUS, an open-source Rust rewrite of WAL-G for PostgreSQL backups, focusing on predictable memory usage and WAL-G compatibility.

0 favorites 0 likes
#memory-efficiency

SHAPE: Coalition-Aware Expert Pruning for Sparse Mixture-of-Experts LLMs

arXiv cs.LG ↗ · 2026-06-10 Cached

SHAPE proposes a coalition-aware expert pruning framework for sparse MoE LLMs that uses Shapley-style attribution over routing traces to identify essential experts, achieving competitive accuracy under 20-40% pruning and reducing GPU memory footprint.

0 favorites 0 likes
#memory-efficiency

@HuggingPapers: Microsoft Research introduces Mirage Latent spatial memory stores 3D scenes directly as latent tokens, skipping the cos…

X AI KOLs Following ↗ · 2026-06-09 Cached

Microsoft Research introduces Mirage, a latent spatial memory that stores 3D scenes as latent tokens, achieving up to 10.57x faster video generation and 55x lower memory use with state-of-the-art consistency.

0 favorites 0 likes
#memory-efficiency

FlashMemory-DeepSeek-V4: Lightning Index Ultra-Long Context via Lookahead Sparse Attention

Hugging Face Daily Papers ↗ · 2026-06-08 Cached

Proposes Lookahead Sparse Attention with a Neural Memory Indexer on DeepSeek-V4, reducing GPU memory usage to ~13.5% of full-context baseline while maintaining or slightly improving accuracy.

0 favorites 0 likes
#memory-efficiency

proveKV – Honest 36× lossless (vs f32, 18x vs fp16) KV‑cache compression for LLMs (zero PPL regression)

Reddit r/LocalLLaMA ↗ · 2026-06-05

An open-source repo, proveKV, demonstrates a reproducible KV-cache compression technique achieving 36x lossless (vs f32) and 68x lossy memory reduction on SmolLM2-1.7B with zero PPL regression, including Rust examples and an audit pipeline.

0 favorites 0 likes
#memory-efficiency

dMoE: dLLMs with Learnable Block Experts

Hugging Face Daily Papers ↗ · 2026-05-29 Cached

This paper proposes dMoE, a block-level mixture-of-experts framework for diffusion large language models that aggregates token-level expert distributions into block-level routing, reducing activated experts and memory usage while maintaining performance.

0 favorites 0 likes
#memory-efficiency

For over a decade, we've accepted that end-to-end backprop is the only way to train deep networks (1 minute read)

TLDR AI ↗ · 2026-05-29 Cached

Sakana AI presents DiffusionBlocks, a method that trains neural networks block-wise by interpreting forward passes as diffusion denoising, significantly reducing memory requirements compared to traditional end-to-end backpropagation.

0 favorites 0 likes
#memory-efficiency

Reparametrizing Shampoo and SOAP for Subspace Basis Updates and BFloat16 Storage

arXiv cs.LG ↗ · 2026-05-27 Cached

This paper proposes a reparametrization of the preconditioner in Shampoo-based optimization methods (like KL-Shampoo and SOAP) to support BFloat16 storage and reduce computational overhead by updating only part of the basis via QR decomposition in a subspace, making these methods more memory- and time-efficient.

0 favorites 0 likes
#memory-efficiency

Quantized Keys Steal Attention: Bias Correction for KV-Cache Compression in Video Diffusion

arXiv cs.LG ↗ · 2026-05-27 Cached

This paper identifies a bias in attention weights caused by quantizing keys in KV-cache compression for chunk-wise autoregressive video diffusion, and proposes a per-attention-score correction that removes the bias with negligible overhead, recovering near-BF16 video quality at INT2 quantization.

0 favorites 0 likes
#memory-efficiency

CONF-KV: Confidence-Aware KV Cache Eviction with Mixed-Precision Storage for Long-Horizon LLM

Hugging Face Daily Papers ↗ · 2026-05-24 Cached

CONF-KV is a KV-cache management system that uses model uncertainty to dynamically adjust cache retention, improving memory efficiency for long-context LLM inference while maintaining accuracy within 1.5-2.1 perplexity points.

0 favorites 0 likes
#memory-efficiency

Self-Pruned Key-Value Attention: Learning When to Write by Predicting Future Utility

arXiv cs.LG ↗ · 2026-05-15 Cached

Introduces Self-Pruned Key-Value Attention (SP-KV), a mechanism that learns to predict future utility of key-value pairs to dynamically prune the KV cache, reducing memory usage and decoding speed by 3-10x with minimal performance degradation. The model and utility predictor are trained end-to-end using next-token prediction.

0 favorites 0 likes
#memory-efficiency

When Does Value-Aware KV Eviction Help? A Fixed-Contract Diagnostic for Non-Monotone Cache Compression

arXiv cs.LG ↗ · 2026-05-12 Cached

This paper introduces a fixed-contract diagnostic tool to analyze why KV cache compression methods succeed or fail in long-context LLM inference. It identifies three failure modes—missing evidence, scoring irrelevant tokens, and breaking related evidence—and evaluates them on LongBench and NeedleBench.

0 favorites 0 likes
#memory-efficiency

ReST-KV: Robust KV Cache Eviction with Layer-wise Output Reconstruction and Spatial-Temporal Smoothing

arXiv cs.CL ↗ · 2026-05-12 Cached

This paper introduces ReST-KV, a novel method for robust KV cache eviction in large language models that uses layer-wise output reconstruction and spatial-temporal smoothing to improve efficiency. The method significantly reduces decoding latency and outperforms state-of-the-art baselines on long-context benchmarks like LongBench and RULER.

0 favorites 0 likes
#memory-efficiency

How to Compress KV Cache in RL Post-Training? Shadow Mask Distillation for Memory-Efficient Alignment

arXiv cs.LG ↗ · 2026-05-11 Cached

This paper proposes Shadow Mask Distillation (SMD) to solve the off-policy bias caused by KV cache compression during reinforcement learning post-training for large language models. It introduces a mechanism that ensures on-policy alignment and improves memory efficiency for long-context reasoning tasks.

0 favorites 0 likes
#memory-efficiency

OjaKV: Context-Aware Online Low-Rank KV Cache Compression

arXiv cs.CL ↗ · 2026-04-20 Cached

OjaKV introduces a context-aware online low-rank KV cache compression framework that uses hybrid storage and Oja's algorithm for incremental subspace adaptation to reduce GPU memory bottlenecks in long-context LLM inference without model fine-tuning.

0 favorites 0 likes
#memory-efficiency

Ulysses Sequence Parallelism: Training with Million-Token Contexts

Hugging Face Blog ↗ · 2026-03-09 Cached

Ulysses Sequence Parallelism is a technique for training LLMs with million-token contexts by distributing sequence chunks across GPUs, reducing memory requirements and enabling efficient long-context training. It integrates with HuggingFace Accelerate, Transformers Trainer, and TRL, with support for Flash Attention and DeepSpeed ZeRO.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback