Tag
Merged fixes for quantized KV cache into the DeepSeek V4 branch of llama.cpp, with benchmark perplexity results for f16, q8_0, and q4_0 cache types.