Just to be clear, owning the hardware isn’t easy either!

Reddit r/singularity News

Summary

A company shares the practical difficulties of purchasing and maintaining an HGX B300 for AI inference, including high cost (€1.1M), power and space requirements, and limited availability, with throughput estimates for GLM 5.2.

My company is currently trying to buy one HGX B300 for glm 5.2. You need around €1.1m just to buy itthat’s the offer we got a month ago. Then you need a place for it and at least 0.25–0.5 FTE for maintenance. It will probably be useful for 2–5 years, but won’t even be the newest generation by the end of this year One HGX B300 gives you around 7k tokens/sec, while one active user needs around 40–50 tokens/sec for glm 5.2 Realistically, that’s around 20 concurrent coding users or 120–180 RAG users. And there aren’t many B300s available, so even with the money, you might not get one!
Original Article

Similar Articles

Running GLM5.2 on budget hardware < $2500.

Reddit r/LocalLLaMA

A guide showing how to build a system under $2500 using used server components to run GLM5.2 and other large AI models locally, with trade-offs in speed.

"Hardware is the only moat" - Should we buy new hardware now or wait?

Reddit r/LocalLLaMA

The article discusses the growing importance of hardware as a competitive advantage in AI, noting that leading labs are prioritizing product competitiveness and compute scale over pure AGI research. It highlights the resulting strain on consumer GPU availability and the increasing costs for hardware upgrades.

Buying AI accelerators/GPUs in China...

Reddit r/LocalLLaMA

A user asks about buying Chinese AI accelerators/GPUs for inference, specifically looking for Huawei alternatives to Nvidia, with support for vLLM or Llama.cpp.