Added an old 2070 Super to my rig and I can't go back...worse, now I need more

Reddit r/LocalLLaMA News

Summary

A user shares their experience of adding an old NVIDIA 2070 Super GPU to their rig for extra VRAM, enabling them to run larger LLMs like Qwen3.6-27B at high quantization and context size with good performance, and now considering upgrading to a 3090 for even more VRAM.

Context: I built a new system last year November before everything went to shit. I spent like 5k for a 5090, 9800X3D and 96GB RAM. Recently (last 2-3 months) I'm heavily working on my local setup. Ditched Windows, went Ubuntu > Manjaro > CachyOS (now) and I'm basically building llama.cpp everyday now running tests to find optimal model quantizations, context sizes, best agent cli + harness, etc...most of you know the drill. Now: I finally got around and took my old PC apart. I saw the 2070, dusted it off and put in my new PC (just out of curiousity). LET ME TELL YOU: I was not ready for what 8GB of additional VRAM does to a mf. I can suddenly run Qwen3.6-27B at Q8_0 with a context of 144k (q8_0 as well) and with MTP and I still generate 40-70tk/s. It's addicting! Now I'm looking at offers online for 5070tis and 3090s (because they are in the same ball park prize wise). I mean it's going to be the 3090 eventually, because I can't just pass on 8GB of VRAM but again I wasn't ready for this. Even a 2070 Super brings so much value if you have it laying around. This experience was eye opening in terms of: acceptable performance + bigger VRAM > amazing performance + smaller VRAM
Original Article

Similar Articles

Best models in 3x3090 (72GB VRAM) in Q2 2026?

Reddit r/LocalLLaMA

A user shares their experience running large LLMs on a 3x3090 (72GB VRAM) setup in Q2 2026, recommending models like GPT-OSS 120b, Qwen3.5 122b, and GLM Air 4.5 106B, and asking for newer alternatives.

Best Settings for 48GB VRAM + Qwen 3.6 27B

Reddit r/LocalLLaMA

A user shares optimized settings for running Qwen3.6 27B (Q8_0) on a dual GPU setup (RTX 4090 + RTX 3090) with llama.cpp, achieving 75-100 t/s and 1500 pp with 250k context.