llama-cpp-comparison

Tag

Cards List
#llama-cpp-comparison

Ninfer and a 5090 with 3.8 27B is making me cry tears of joy it's so good.

Reddit r/LocalLLaMA · 2d ago

A user reports achieving high token throughput with the Ninfer tool on an NVIDIA RTX 5090 GPU using a Qwen 3.8B model, significantly outperforming llama.cpp.

0 favorites 0 likes
← Back to home

Submit Feedback