Tag
Showcases the performance of the Qwen3.8-27B AI model running on a 12GB RTX 3080 Ti GPU, achieving 47 tokens per second at 128K context using llama.cpp with specific quantization settings.
The RTX 3080 20GB variant is reportedly available for $438, considered a good deal for a high-end GPU.