Tag
A demonstration of Qwen 3.8 27B Q8 model's performance on dual NVIDIA 3090 GPUs, creating a Minecraft clone named RetroCraft in a single prompt with detailed token processing and generation speeds.
After two months of local LLM testing, the author finds that the combination of gemma-4-12B-it-QAT and MTP assistance performs best in speed and usability, with hardware i7-13700 + 64GB RAM + RTX 4070.
A user shares their experience setting up a dual-GPU local AI lab with RTX 4080 Super and 5060 Ti, running Qwen 3.6 models via llama.cpp and llama-swap to reduce API costs and enable unrestricted experimentation.