Tag
Demonstrates achieving 20GB VRAM and 448GB/s bandwidth for around $100 using two NVIDIA P102-100 cards, running a llama.cpp server with a Qwen model and supporting 3 concurrent users with large context.