Tag
WonderBox is a native macOS system toolbox that provides system monitoring, memory optimization, cache cleaning, and other utilities to help users manage their Mac efficiently.
A user shares their experience switching to the QWEN3.6-27B-MLX-8bit model for local AI tasks, finding it performs comparably to larger models while saving significant RAM on a Mac Studio, improving workflow stability.
This paper introduces methods to bound four key memory peaks in large Mixture-of-Experts training, enabling training at 1M context length with fixed GPU memory and up to 10.4× throughput improvement over baselines.
The paper introduces PAGE, a partition-aware gated KV-cache eviction method that uses a scalar metric to predict input classes and apply eviction only when safe, reducing accuracy degradation in large language models.
TierKV proposes a predictive multi-tier KV caching framework to optimize memory usage and throughput for long-context LLMs on mobile devices, achieving significant performance improvements with minimal accuracy degradation.
Cloudflare optimized the memory usage of their Pingora-based load-balancing service by refining the pingora-ketama consistent hashing library in Rust, reclaiming over 100TB of RAM globally.
Memorable introduces a method to optimize memory in AI agents using embeddings instead of tokens, enabling procedural memory that persists across multiple runs.
Edge0 is a streaming MoE inference engine that uses trained routing prediction to serve 35B Mixture-of-Experts models from SSD, achieving near-fp16 performance on consumer hardware with low memory usage.
This paper introduces techniques to manage memory peaks in training large Mixture-of-Experts models with long context lengths, including Pipelined LLEP, Ring-DTP, SCO, and OffloadStreamAdamW, which enable fixed GPU working sets and improve throughput up to 10.4x.
BeaconKV is a training-free method that uses beacon queries to compress key-value cache for efficient inference in Large Reasoning Models, reducing memory usage and improving throughput without sacrificing accuracy.
HeadWiseKV is a training-free framework that compresses KV caches in hybrid long-context language models, reducing GPU memory usage and extending context lengths while maintaining quality.
The paper proposes Random Attention, a KV cache eviction method that uses random selection instead of scoring, matching selective methods while improving throughput in reasoning tasks.
A video explaining memory optimization techniques used in the Superlogical server, including in-memory compression, dynamic thread management, and fast binary snapshots.
slotstream is an open-source tool that enables running large AI models like Qwen3.8-Flash-Next on Macs with limited memory by streaming model weights from SSD, achieving approximately 12 tokens per second on a 48GB Mac.
FlashAccel is a co-designed system that integrates High-Bandwidth Flash into GPUs to enhance LLM inference, achieving improved throughput and energy efficiency by mitigating access latency and optimizing bandwidth utilization.
A new checkpoint for Qwen3.8 Flash-Next offloads the ngram lookup table to SSD, reducing RAM usage by 48 GB while maintaining inference speed in SGLang.
Zod 4.5 implements method memoization to drastically reduce memory usage, achieving up to 9.8x less heap retention than previous versions.
Sliding window attention with sinks outperforms post-trained linear attention on long-context tasks without retraining, offering a cheaper and more reliable inference solution for LLMs.
Cloudflare optimized their 1.1.1.1 DNS cache to reduce memory usage by over 50%, saving 100 terabytes of memory while improving performance through refined Rust data structures.
This paper analyzes KV-cache eviction strategies in sparse-attention models, showing that selecting largest weights is near-optimal and that published margins come from memory and query information, with ContourKV achieving strong performance.