I built a server with 768GB VRAM for frontier, but all new frontier open source models are likely to be two trillion or above now, including next GLM 6, am I cooked?

Reddit r/LocalLLaMA News

Summary

An individual discusses building a server with 768GB VRAM for running frontier AI models but is concerned that new open-source models like GLM6 are becoming too large, prompting consideration of downsizing to smaller flash models.

This epyc server I am using twelve cards with 64 GB memory, plus 256GB ram. Looking at the most capable models in open source, GLM 5.3 seems to be the only option, but with Astra releasing it will likely be fairly behind. GLM6 looks like it will be at least double in size, maybe even triple. Qwen-max and Kimmi are already way too big to even consider. Even the deepseek V4 Pro is too big. Should I just give up on this frontier dream sell the excess GPUs and settle For flash models with far fewer GPUs and a reasonable cost. Note: I'm not using it for any business. I was hoping to build a new business with this, but it can probably be done with much more effort with a flash model as well. Edit: I don't want to go below 4-bit quants because then the models start making obvious mistakes. So I'm talking about a min/max of 4-bit Okay, this post really blew up. I wasn't expecting so much interest or comments just attacking me. Was really just expecting to have a calm discussion about future SOTA model sizes.
Original Article

Similar Articles

12GB VRAM gang, what's our plan?

Reddit r/LocalLLaMA

Discussion about running LLMs on 12GB VRAM, noting current focus on dense models like Muse Glimmer 30B and Qwen 3.8 27B, and questioning whether upgrading to 24GB VRAM is needed.

Best models in 3x3090 (72GB VRAM) in Q2 2026?

Reddit r/LocalLLaMA

A user shares their experience running large LLMs on a 3x3090 (72GB VRAM) setup in Q2 2026, recommending models like GPT-OSS 120b, Qwen3.5 122b, and GLM Air 4.5 106B, and asking for newer alternatives.

How many people have 24gb over gpu here?

Reddit r/LocalLLaMA

The author discusses the low adoption of the qwen 3.8 27b model based on download counts and estimates that very few users have the high-VRAM GPUs needed for productive local LLM development.