Tag
A user reports achieving high token throughput with the Ninfer tool on an NVIDIA RTX 5090 GPU using a Qwen 3.8B model, significantly outperforming llama.cpp.
The author discusses the high cost of the upcoming NVIDIA RTX 5090 GPU, considering an Apple Mac Studio with M5 Ultra as an alternative and expressing concern over the situation.