memory-optimization

Tag

Cards List
#memory-optimization

@QingQ77: WonderBox is a system toolbox that runs entirely natively on Mac, centrally providing system monitoring, memory optimiz…

X AI KOLs Timeline ↗ · 2d ago Cached

WonderBox is a native macOS system toolbox that provides system monitoring, memory optimization, cache cleaning, and other utilities to help users manage their Mac efficiently.

0 favorites 0 likes
#memory-optimization

QWEN3.6-27B-MLX-8bit (29.5GIG) Is Excellent

Reddit r/openclaw ↗ · 3d ago

A user shares their experience switching to the QWEN3.6-27B-MLX-8bit model for local AI tasks, finding it performs comparably to larger models while saving significant RAM on a Mac Studio, improving workflow stability.

0 favorites 0 likes
#memory-optimization

Keeping Large MoE Training Within Fixed GPU Memory (20 minute read)

TLDR AI ↗ · 5d ago Cached

This paper introduces methods to bound four key memory peaks in large Mixture-of-Experts training, enabling training at 1M context length with fixed GPU memory and up to 10.4× throughput improvement over baselines.

0 favorites 0 likes
#memory-optimization

PAGE: Partition-Aware Gated KV-Cache Eviction

arXiv cs.LG ↗ · 6d ago Cached

The paper introduces PAGE, a partition-aware gated KV-cache eviction method that uses a scalar metric to predict input classes and apply eviction only when safe, reducing accuracy degradation in large language models.

0 favorites 0 likes
#memory-optimization

TierKV: Long-Context On-Device LLMs via Predictive Multi-Tier KV Caching

arXiv cs.LG ↗ · 2026-09-21 Cached

TierKV proposes a predictive multi-tier KV caching framework to optimize memory usage and throughput for long-context LLMs on mobile devices, achieving significant performance improvements with minimal accuracy degradation.

0 favorites 0 likes
#memory-optimization

Saving another 100TB of RAM

Hacker News Top ↗ · 2026-09-18 Cached

Cloudflare optimized the memory usage of their Pingora-based load-balancing service by refining the pingora-ketama consistent hashing library in Rust, reclaiming over 100TB of RAM globally.

0 favorites 0 likes
#memory-optimization

@garrytan: Memorable found a way to optimize memory with embeddings instead of more tokens which is a powerful new way to do memory

X AI KOLs Timeline ↗ · 2026-09-17 Cached

Memorable introduces a method to optimize memory in AI agents using embeddings instead of tokens, enabling procedural memory that persists across multiple runs.

0 favorites 0 likes
#memory-optimization

The Other Half of the Memory Wall: Serving 35B MoEs from SSD with Trained Routing Prediction

arXiv cs.AI ↗ · 2026-09-17 Cached

Edge0 is a streaming MoE inference engine that uses trained routing prediction to serve 35B Mixture-of-Experts models from SSD, achieving near-fp16 performance on consumer hardware with low memory usage.

0 favorites 0 likes
#memory-optimization

Flattening Every Memory Peak in Long-Context Mixture-of-Experts Training

Hugging Face Daily Papers ↗ · 2026-09-13 Cached

This paper introduces techniques to manage memory peaks in training large Mixture-of-Experts models with long context lengths, including Pipelined LLEP, Ring-DTP, SCO, and OffloadStreamAdamW, which enable fixed GPU working sets and improve throughput up to 10.4x.

0 favorites 0 likes
#memory-optimization

BeaconKV: Key-Value Cache Compression Guided by Beacon Queries for Efficient Large Reasoning Model Inference

Hugging Face Daily Papers ↗ · 2026-09-04 Cached

BeaconKV is a training-free method that uses beacon queries to compress key-value cache for efficient inference in Large Reasoning Models, reducing memory usage and improving throughput without sacrificing accuracy.

0 favorites 0 likes
#memory-optimization

HeadWiseKV: Budgeted Per-Head Cache Residency for Hybrid Long-Context Language Models

arXiv cs.AI ↗ · 2026-09-03 Cached

HeadWiseKV is a training-free framework that compresses KV caches in hybrid long-context language models, reducing GPU memory usage and extending context lengths while maintaining quality.

0 favorites 0 likes
#memory-optimization

Random Attention: Rethinking KV Cache Eviction for Efficient Reasoning

Hugging Face Daily Papers ↗ · 2026-09-03 Cached

The paper proposes Random Attention, a KV cache eviction method that uses random selection instead of scoring, matching selective methods while improving throughput in reasoning tasks.

0 favorites 0 likes
#memory-optimization

@mitchellh: A video explaining a handful of the tricks we use to optimize for memory within the Superlogical server: in-memory comp…

X AI KOLs Timeline ↗ · 2026-09-02 Cached

A video explaining memory optimization techniques used in the Superlogical server, including in-memory compression, dynamic thread management, and fast binary snapshots.

0 favorites 0 likes
#memory-optimization

Show HN: Running 104GB Qwen3.8-Flash-Next on 48GB Mac with at ~12 tok/s

Hacker News Top ↗ · 2026-09-01 Cached

slotstream is an open-source tool that enables running large AI models like Qwen3.8-Flash-Next on Macs with limited memory by streaming model weights from SSD, achieving approximately 12 tokens per second on a 48GB Mac.

0 favorites 0 likes
#memory-optimization

{INTRESTING PAPER BASED ON HBF}2607.10186] FlashAccel: Leveraging High-Bandwidth Flash (HBF) for High-Throughput LLM Inference

Reddit r/LocalLLaMA ↗ · 2026-08-30 Cached

FlashAccel is a co-designed system that integrates High-Bandwidth Flash into GPUs to enhance LLM inference, achieving improved throughput and energy efficiency by mitigating access latency and optimizing bandwidth utilization.

0 favorites 0 likes
#memory-optimization

Qwen 3.8 Flash Next ngram look up table offloaded to SSD and streamed in SGLang

Reddit r/LocalLLaMA ↗ · 2026-08-29 Cached

A new checkpoint for Qwen3.8 Flash-Next offloads the ngram lookup table to SSD, reducing RAM usage by 48 GB while maintaining inference speed in SGLang.

0 favorites 0 likes
#memory-optimization

@southpolesteve: Zod 4.5 comes with a massive drop in memory usage. If you are running in @Cloudflare Workers you should update!

X AI KOLs Timeline ↗ · 2026-08-29 Cached

Zod 4.5 implements method memoization to drastically reduce memory usage, achieving up to 9.8x less heap retention than previous versions.

0 favorites 0 likes
#memory-optimization

Sliding-window beats linear attention

Hugging Face Daily Papers ↗ · 2026-08-28 Cached

Sliding window attention with sinks outperforms post-trained linear attention on long-context tasks without retraining, offering a cheaper and more reliable inference solution for LLMs.

0 favorites 0 likes
#memory-optimization

Saving 100 terabytes of memory by optimizing 1.1.1.1's DNS cache

Hacker News Top ↗ · 2026-08-27 Cached

Cloudflare optimized their 1.1.1.1 DNS cache to reduce memory usage by over 50%, saving 100 terabytes of memory while improving performance through refined Rust data structures.

0 favorites 0 likes
#memory-optimization

Trust the Mass: Forced Weights in KV-Cache Eviction

arXiv cs.LG ↗ · 2026-08-27 Cached

This paper analyzes KV-cache eviction strategies in sparse-attention models, showing that selecting largest weights is near-optimal and that published margins come from memory and query information, with ContourKV achieving strong performance.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback