Tag
A user configured a mismatched pair of Tesla V100 GPUs (16GB and 32GB) into a capable local LLM lab using llama.cpp with tensor split and other optimizations, achieving high prompt and decode speeds with the Qwen3.8 27B model.