Tag
User seeks advice on preventing llama.cpp from offloading KV cache to swap before RAM is fully exhausted, sharing their configuration on an M2 Max with 96GB RAM and a large Qwen model.
nbd-vram is a Linux tool that uses NVIDIA GPU VRAM as swap space via the NBD protocol and CUDA, providing extra memory for systems with soldered RAM and no upgrade path.