I built a server with 768GB VRAM for frontier, but all new frontier open source models are likely to be two trillion or above now, including next GLM 6, am I cooked?
Summary
An individual discusses building a server with 768GB VRAM for running frontier AI models but is concerned that new open-source models like GLM6 are becoming too large, prompting consideration of downsizing to smaller flash models.
Similar Articles
12GB VRAM gang, what's our plan?
Discussion about running LLMs on 12GB VRAM, noting current focus on dense models like Muse Glimmer 30B and Qwen 3.8 27B, and questioning whether upgrading to 24GB VRAM is needed.
Is there a GLM 5.3 Flash Antirez/DS4 GGUF targeted at 192 GB RAM?
The author asks whether a GLM 5.3 Flash Antirez/DS4 GGUF model around 192 GB RAM exists and seeks advice on creating such a model.
High VRAM local coding model — still Qwen 3.6 27B?
The user discusses their experience with Qwen 3.6 27B for local coding tasks and asks for recommendations for larger models (100B+) suitable for systems with 224GB of VRAM.
Best models in 3x3090 (72GB VRAM) in Q2 2026?
A user shares their experience running large LLMs on a 3x3090 (72GB VRAM) setup in Q2 2026, recommending models like GPT-OSS 120b, Qwen3.5 122b, and GLM Air 4.5 106B, and asking for newer alternatives.
How many people have 24gb over gpu here?
The author discusses the low adoption of the qwen 3.8 27b model based on download counts and estimates that very few users have the high-VRAM GPUs needed for productive local LLM development.