Measuring PCIe transfer under dual GPU with pipeline & tensor llama.cpp

Reddit r/LocalLLaMA Tools

Summary

An analysis of PCIe transfer performance when running llama.cpp with dual GPUs using pipeline and tensor parallelism.

No content available
Original Article

Similar Articles

Dual GPU llama.cpp speedup

Reddit r/LocalLLaMA

A fork of llama.cpp fixes the --split-mode tensor issue with quantized KV caches, achieving up to 40% speed improvement on dual GPU setups without quality loss.