cpu-offloading

Tag

Cards List
#cpu-offloading

@QuixiAI: LESSON LEARNED: Always use BF16 kv cache. I was using turboquant. Yeah the VRAM consumption sucks - so another trick is…

X AI KOLs Timeline · 4d ago Cached

The author shares a technical lesson on using BF16 KV cache instead of turboquant for AI model optimization and implements CPU offloading in SlimServe for Qwen models to manage VRAM consumption.

0 favorites 0 likes
#cpu-offloading

[Deepseek-V4-Flash-0731] Full 1M context on a single RTX5090 + DDR5 Desktop Setup with VLLM CPU/Ram Offloading, ~800 tps pp & 15+ tps decode [Agentic Coding]

Reddit r/LocalLLaMA · 2026-08-04

Technical write-up on running DeepSeek-V4-Flash-0731 with full 1M context on a single RTX 5090 + 256GB DDR5 desktop using vLLM with CPU/RAM offloading, achieving ~800 tps prefill and ~15 tps decode.

0 favorites 0 likes
#cpu-offloading

We could really use Qwen3.8 in 27B, 35B, 122B and 397B sizes

Reddit r/LocalLLaMA · 2026-07-27

The author argues that releasing smaller Qwen models (27B, 35B, 122B, 397B) would better serve the local AI community than focusing on trillion-parameter behemoths, which are impractical for most users.

0 favorites 0 likes
#cpu-offloading

DeepSeek V4 Flash (98GB) on 1x 4060ti + CPU got 300% faster this week [ 2->7t/s]

Reddit r/LocalLLaMA · 2026-07-16

DeepSeek V4 Flash (98GB) now runs up to 7 tokens per second on a single RTX 4060 Ti with CPU offloading, a 3x speed improvement over the previous week's 2 t/s.

0 favorites 0 likes
← Back to home

Submit Feedback