4xR9700, 2xMi210 or 4x4080S 32G

Reddit r/LocalLLaMA News

Summary

The user is comparing GPU options like R9700, Mi210, and 4080S to achieve 128GB VRAM for running multiple AI models in parallel, considering factors like cost, performance, and compatibility.

I am trying to get to 128G of VRAM with reasonable compute and bandwidth to run multiple models in parallel. DS4 Flash or GLM 5.3 in hybrid mode with custom checkpoints. But I am a bit confused these days given the million forks, I only know a bit about the NVidia world. So one option is to use modded 4080S which are around 1.7K units of currency, which has the advantage of being CUDA. It has 736 GB/sec memory bandwidth but no NVFP4 support. The RTX Pro 4500 is too expensive at 2.7K. Then comes the R9700 with 640 GB/sec of memory bandwidth at around 1.5K as well. But it is ROCm/Vulkan. It seems it is growing fast? Is there somewhere I can read up to date and trustable comparisons of R9700 with equivalent-ish Intel and Nvidia ones? And then comes the Mi210. PCIe Gen4, HBM2x64G. 3.5K cost, so linearly double. Half the power consumption and 1600 GB/sec of bandwidth. Looks awesome! But it seems the only support it has for newer model formats is simply upscale to BF16 and run with those. Do I read it correctly that it's basically like having a 16G VRAM card with the new NVFP4/MXFP4 models? And a 32G for the FP8 ones? Or is this the gem i should buy instead? (No, I will not buy more Sparks).
Original Article

Similar Articles

Best models in 3x3090 (72GB VRAM) in Q2 2026?

Reddit r/LocalLLaMA

A user shares their experience running large LLMs on a 3x3090 (72GB VRAM) setup in Q2 2026, recommending models like GPT-OSS 120b, Qwen3.5 122b, and GLM Air 4.5 106B, and asking for newer alternatives.