1x32GB V100 vs 2x16GB V100 vs 5060ti 16GB for QWEN 3.8

Reddit r/LocalLLaMA News

Summary

A user is considering upgrading from a 5060ti 16GB to either 1x32GB V100 or 2x16GB V100 for better performance with the Qwen 3.8 model using llama.cpp, and asking for other options in a similar price range.

Hi All, I am currently contemplating an upgrade from my 5060ti 16gb. I am getting ~40t/s with 130k context on Qwen 3.8 IQ3_S HF quant. I am running llama.cpp on linux. Objective is to increase context and use a better quant and also free up 5060 for other tasks. The options I am considering are 1x32GB V100 and 2x16GB V100. Theoretically, 2x16GB should be superior in terms of performance to 5060 and 1x32gb due to higher memory bandwidth. One issue I have to deal with is that I am limited in terms of CPU to GPU comms - I only have 2x x4 lines available. Any other good options in the same price range?
Original Article

Similar Articles

TQwen 3.8 flash next ud1s on 6gb vram and 16 gb system ram

Reddit r/LocalLLaMA

A user shares their experience running the Qwen 3.8 flash next model on a system with 6GB VRAM and 16GB RAM using llama.cpp, achieving 6-7 tokens per second with 1-bit quantization, and asks for recommendations on quantization variants.