GPU guide (GB per dollar, bandwidth)
Summary
The article provides a guide comparing GPUs based on cost per gigabyte and bandwidth to help optimize script performance.
Similar Articles
I compared all specs of the major GPUs/machines that are being used here, because bandwidth is not everything. Some of ya'll need a reality check.
The author compares various GPUs for LLM inference, critiquing common benchmarks and emphasizing the importance of prefill performance over generation speed, offering recommendations for different budgets and use cases.
Computable GPU Index (CGI)
Computable GPU Index (CGI) is an open-source price index that calculates the USD price per GPU-hour from the rental rates of various providers. It provides a mathematically robust and reproducible method for benchmarking GPU compute costs.
Performance per dollar is getting faster and cheaper
Wafer demonstrates that AMD MI355X GPUs offer competitive inference performance for frontier models like GLM5.2 at significantly lower cost than NVIDIA Blackwell, achieving 80% of B200 throughput at under half the price, using MXFP4 quantization and sglang.
A GPU-Hour Isn't a Commodity If You Need Four of Them (6 minute read)
An analysis of GPU rental prices on Vast.ai reveals that pricing for multi-GPU configurations deviates significantly from per-GPU-hour headline rates due to supply constraints, with larger clusters often being unavailable or more expensive.
@akshay_pachaar: GPU architecture, clearly explained. The usual assumption is that a faster GPU means more compute, so a chip rated for …
The article clarifies that GPU performance in AI inference is limited by memory bandwidth rather than compute power, using the NVIDIA H100 as an example to explain GPU architecture and its effect on token generation rates.