memory-efficiency

Tag

Cards List
#memory-efficiency

GEM-KMeans: Memory-Efficient and Accurate Clustering on Massive Scale with GPU Optimization

arXiv cs.LG ↗ · 11h ago Cached

GEM-KMeans introduces a spectrally normalized, memory-efficient NLR formulation for K-means clustering that fuses updates into a single matrix-multiplication epilogue, reducing GPU HBM storage to one factor while maintaining clustering accuracy at massive scale.

0 favorites 0 likes
#memory-efficiency

Pruned CTC for Memory-Efficient Large-Vocabulary ASR Training

Hugging Face Daily Papers ↗ · 3d ago Cached

This paper introduces Pruned CTC, a method that reduces memory usage in Connectionist Temporal Classification training for large-vocabulary automatic speech recognition by restricting alignment computation, proving equivalence to full-vocabulary CTC and demonstrating strong performance with LLM adaptations in offline and streaming settings.

0 favorites 0 likes
#memory-efficiency

@hooshaaii: LLMs crash when their KV cache exceeds memory limits. The paper "Efficient Streaming Language Models with Attention Sin…

X AI KOLs Timeline ↗ · 6d ago Cached

The paper introduces StreamingLLM, a framework that uses attention sinks to enable large language models to handle infinite sequence lengths without fine-tuning, improving efficiency in streaming applications.

0 favorites 0 likes
#memory-efficiency

Flash-dLLM: IO-Aware KV Caching and Parallel Decoding for Fast, Memory-Efficient Diffusion LLMs

Hugging Face Daily Papers ↗ · 2026-09-22 Cached

Flash-dLLM is a training-free inference acceleration framework for diffusion LLMs that uses IO-aware KV caching and parallel decoding to achieve significant speedups and memory efficiency improvements.

0 favorites 0 likes
#memory-efficiency

SGD-KV: Summarization Guided KV Cache Compression

arXiv cs.CL ↗ · 2026-09-04 Cached

SGD-KV is a framework that uses summarization to guide KV cache compression in large language models, reducing memory usage by up to 75% for contexts up to 1M tokens while achieving state-of-the-art performance on long-context benchmarks.

0 favorites 0 likes
#memory-efficiency

PuzzleKV: Page-Wise Low-Rank Decomposition for KV Cache Compression

arXiv cs.LG ↗ · 2026-08-26 Cached

PuzzleKV is a training-free method for compressing key-value cache in large language models using page-wise low-rank decomposition, achieving over 96% performance with approximately 60% storage.

0 favorites 0 likes
#memory-efficiency

Prefix Sliding for efficient test-time scaling

Hugging Face Daily Papers ↗ · 2026-08-26 Cached

Prefix Sliding reduces memory costs during long reasoning by discarding unimportant intermediate tokens, enabling efficient test-time scaling without retraining, achieving up to 3x speedup in existing models.

0 favorites 0 likes
#memory-efficiency

Gated Recurrent Transformers: Expressive Depth through Recurrent Modulation in Transformers

Hugging Face Daily Papers ↗ · 2026-08-25 Cached

The paper introduces Gated Recurrent Transformer, a model that reuses a shared core across depth with adaptive update gates, achieving comparable or better quality than deeper models with fewer parameters and lower memory.

0 favorites 0 likes
#memory-efficiency

Q-Interference: Memory-Efficient Phase-Aware Quantum-Inspired Attention

arXiv cs.CL ↗ · 2026-08-19 Cached

This paper proposes Q-Interference, a memory-efficient quantum-inspired attention mechanism for GPT models that uses phase-aware scoring and an exact trigonometric factorization to improve token interaction without increasing memory overhead.

0 favorites 0 likes
#memory-efficiency

p-Spin Glass Network Efficient Single-Batch Continual Learning

arXiv cs.LG ↗ · 2026-08-18 Cached

Introduces the p-Spin Glass Network, a novel architecture for sequence models that achieves memory efficiency, sample efficiency, and single-batch stability, enabling continual learning and edge AI applications.

0 favorites 0 likes
#memory-efficiency

Forward Pass Domain Adaptation (Without Cross-Layer Backpropagation)

arXiv cs.LG ↗ · 2026-08-18 Cached

Forward-Pass-Only MLP training (FPO) adapts large language models without backpropagation, achieving 2.7–3.2× higher throughput and 40% less peak memory while maintaining benchmark performance.

0 favorites 0 likes
#memory-efficiency

RecurrentGPT: Expressive Depth through Recurrent Modulation in Transformers

arXiv cs.CL ↗ · 2026-08-18 Cached

RecurrentGPT introduces a recurrent depth transformer that uses gated modulation to iteratively reuse shared layers, achieving competitive accuracy with fewer parameters and improved memory efficiency compared to standard transformers.

0 favorites 0 likes
#memory-efficiency

When Models Learn (4 minute read)

TLDR AI ↗ · 2026-08-18 Cached

This article explains test-time training, where AI models adapt during inference to improve personalization and reduce memory usage, but at the cost of increased per-user compute. It discusses implications for serving models at scale, balancing long context and user concurrency.

0 favorites 0 likes
#memory-efficiency

Predicting When Random Low-Dimensional Reparameterizations Train Neural Networks

arXiv cs.LG ↗ · 2026-08-14 Cached

This paper analyzes when random low-dimensional reparameterizations can train neural networks, deriving an orientation-resolved master formula for the random-slice residual and introducing RaMaN, a scalable framework that predicts required latent dimensions while dramatically reducing memory costs.

0 favorites 0 likes
#memory-efficiency

Efficient Reinforcement Learning for Long-Horizon Tool-Use Agentic Tasks

arXiv cs.LG ↗ · 2026-08-12 Cached

This paper introduces SinkFlex-RL, a modular training system for memory-feasible reinforcement learning in long-horizon tool-use agentic tasks. It combines a Gymnasium-compatible environment wrapper, GRPO-based policy optimization, and a sink-aware FlexAttention path, reducing peak VRAM by 19.7% at 4096 tokens and enabling 8192-token runs where eager attention runs out of memory.

0 favorites 0 likes
#memory-efficiency

@PrajwalTomar_: STOP. Before you add another AI agent, read this. People are now running 20 coding agents in parallel. TWENTY. And tool…

X AI KOLs Timeline ↗ · 2026-08-11 Cached

The tweet argues that running too many AI coding agents in parallel degrades codebases and advocates a structured setup with a few specialized agents. It also quotes the launch of Jcode, an open-source agent claiming 20x memory efficiency.

0 favorites 0 likes
#memory-efficiency

QEvict: Recoverable Quantized KV Eviction for Attention-Drift-Robust Long-Context Decoding

arXiv cs.LG ↗ · 2026-08-07 Cached

This paper introduces QEvict, a KV-cache management scheme for LLMs that uses recoverable quantized eviction to handle attention drift during long-context decoding, improving memory efficiency while preserving important historical context.

0 favorites 0 likes
#memory-efficiency

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models

arXiv cs.AI ↗ · 2026-07-28 Cached

MM-ShiftKV is a training-free method that improves KV cache selection for multimodal LLMs by approximating decoding-time query behavior during prefilling, reducing memory footprint while preserving performance.

0 favorites 0 likes
#memory-efficiency

@liambraus: CLAUDE CODE TAKES 3.4 SECONDS TO START UP. THIS ONE DOES IT IN 14 MILLISECONDS. A single developer just built their own…

X AI KOLs Timeline ↗ · 2026-07-24 Cached

A single developer built a code agent harness in Rust that starts up 245 times faster than Claude Code (14 ms vs 3.4 s) and uses up to 20x less RAM, achieving 10.4k stars on GitHub.

0 favorites 0 likes
#memory-efficiency

Towards Efficient Large Language Model Serving: A Survey on System-Aware KV Cache Optimization

arXiv cs.LG ↗ · 2026-07-10 Cached

A survey that systematically reviews system-aware KV cache optimization techniques for efficient large language model serving, organizing existing work into execution/scheduling, placement/migration, and representation/retention dimensions.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback