@iluciddreaming: Played with local LLMs for two months. Extensively tested various open-source models using Windows 11 + llama.cpp + llama-swap. Here is my final report card: Hardware: i7-13700 + 64GB RAM + RTX 4070. The best combination currently is gemm…

X AI KOLs Timeline News

Summary

After two months of local LLM testing, the author finds that the combination of gemma-4-12B-it-QAT and MTP assistance performs best in speed and usability, with hardware i7-13700 + 64GB RAM + RTX 4070.

I've been playing with local LLMs for two months. Using Windows 11 + llama.cpp + llama-swap, I extensively tested various open-source models. Here is my final report card: Hardware: i7-13700 + 64GB RAM + RTX 4070 The best combination currently is gemma-4-12B-it-qat-UD-Q4_K_XL + MTP assistance. I've also run other models, but this one currently offers the best balance of speed and practical usability. Next, I'm looking forward to DiffusionGemma coming to llama.cpp, then I'll continue experimenting.
Original Article
View Cached Full Text

Cached at: 06/15/26, 05:07 PM

I’ve been playing with local LLMs for two months.

Using Windows 11 + llama.cpp + llama-swap, I extensively tested various open-source models. Here’s my final report card:

Hardware: i7-13700 + 64GB RAM + RTX 4070

The best-performing combination so far is gemma-4-12B-it-qat-UD-Q4_K_XL + MTP support.

I’ve also run other models, but this one offers the best balance of speed and practical usability.

Next, I’m looking forward to DiffusionGemma being integrated into llama.cpp so I can continue experimenting.

Similar Articles