Cached at:
07/24/26, 05:01 PM
# A GPU-Hour Isn't a Commodity If You Need Four of Them
Source: [https://davefriedman.substack.com/p/a-gpu-hour-isnt-a-commodity-if-you](https://davefriedman.substack.com/p/a-gpu-hour-isnt-a-commodity-if-you)
*Please see relevant disclosures[here](https://davefriedman.substack.com/about)\.*
I recently surveyed[1](https://davefriedman.substack.com/p/a-gpu-hour-isnt-a-commodity-if-you#footnote-1)GPU rental prices on Vast\.ai, the compute rental marketplace\. An H200 rented for about $3\.93 per GPU\-hour\. Without knowing anything else, it looks pretty straightforward\. You pick the chip, you multiply the hourly rate by the number of GPUs you need and the number of hours you need, and you get an invoice\.
But if you want four identical H200s in the same machine, half the qualifying supply disappears, and the remaining supply costs 4% more\. If you want eight, you’re out of luck: there’s nothing in the sample at any price\. The rental rate and the price of usable compute are different numbers\.
In other words, compute pricing coverage runs on headline rates, but many workloads don’t consume them that way\. The headlines say that an H100 costs one amount, and a B200 a different amount\. But training, fine\-tuning, and high\-throughput inference need multiple identical GPUs in the same machine or cluster, with adequate interconnect and a long enough availability window\. Eight scattered GPU\-hours don’t substitute for one hour on eight co\-located GPUs\.
To size the gap between headline numbers and prices for real workloads, I pulled on\-demand offers from Vast\.ai and kept the ones that were verified, currently rentable, available for at least seven days, and attached to a host with a reliability score of 99% or better\.
The sample that resulted from this filtering covered five accelerator models: A100 SXM4, H100 SXM, H200, B200 and L40S\. Multiple offers often represent different bundle sizes on the same box, so I deduplicated by physical machine\. The result was 37 machines and 93 listed GPUs\.
Then I asked how much supply could satisfy a requirement for one, two, four, or eight identical GPUs, and what the best available price was at each size\. The price measure is the median of up to the three cheapest eligible machines\. This limits the pull of a single unusually cheap listing\.
[](https://substackcdn.com/image/fetch/$s_!bEtu!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15c78f66-2154-49cd-a6d8-bae844c1da38_1800x1240.png)*Diagram generated by GPT 5\.6 Sol*
An interesting result of this analysis is that larger clusters don’t necessarily cost more\. Supply of large clusters thins as the requested cluster grows, usually with no matching move in the quoted price\. H200 shows this most cleanly\. Going from one GPU to two left 92% of listed physical inventory eligible at essentially the same price\. Going to four cut eligible inventory to 50% while the price measure rose 4%, from $3\.93 to $4\.08 per GPU\-hour\. At eight GPUs, nothing qualified\.
Supply of H100 and B200 were thinner still\. Qualifying H100 supply supported a two\-GPU request but not a four\-GPU one\. The B200 sample held two physical machines, both topping out at two GPUs, so four\- and eight\-GPU availability was zero\. A100 behaved like you would expect: price increased as supply declined\. Eight\-GPU machines existed, and the price measure climbed from $0\.60 per GPU\-hour at the one\-GPU baseline to $1\.29, a contiguity premium of 114%\. L40S cut both ways\. Only 35% of its listed inventory sat in machines big enough for an eight\-GPU request, but the one machine that qualified was cheap\. This is a thin cross\-sectional order book, not a smooth supply curve\.
Usually, scarcity affects prices\. Think about oil: if demand rises while supply stays constant, prices rise\. But this is not what seems to happen with compute\. Compute rations through configuration and availability instead\. A marketplace can hold hundreds of GPU\-hours in aggregate while a buyer still can’t assemble them into a working cluster\. The binding constraints are hardware generation, co\-location, networking, reliability, reservation duration, CPUs, storage and geography\.
But when there are no suitable configurations, scarcity never shows up as a higher transaction price\. The trade doesn’t happen\. So an index built from one\- and two\-GPU rental prices makes compute look abundant and stable\. Yet at the same time a company scrambling to find a large cluster can’t find one at any price\.
This explains why large buyers rely on bilateral capacity agreements\. Those contracts secure a specific configuration at a specific time\. The excess capacity ends up on marketplaces like Vast\.ai\.
We can assume that the first compute futures contracts will settle against standardized GPU\-hour benchmarks\. This makes sense, since derivatives need simple definitions and enough observations to support a credible index\. However, a company that buys H100 futures to hedge a future four\- or eight\-GPU requirement carries real basis risk\. The futures will hedge a rise in the broad H100 price and do nothing about the lack of co\-located capacity\. The hedge pays but the workload doesn’t run\.
Energy markets manage this residual risk as location basis and the spread between firm and interruptible supply\. Similarly, the compute markets likely will build a version around topology\. Think generic GPU\-hours at the center, with premiums or separate contracts for cluster size, interconnect quality, region and availability guarantees\. And these cluster\-hours and capacity options may well end up trading more actively than the base GPU\-hour\.
Vast\.ai is a single marketplace\. Its listings are advertised offers, not completed transactions, and its visible supply is not global installed capacity\. Several of the larger cluster estimates rest on a handful of machines, so treat the precise premiums with caution\.
The availability pattern is harder to dismiss\. It held when I varied the minimum reliability threshold between 98% and 99\.5%\. B200 still had no four\-GPU offer\. H200 still had no eight\-GPU offer\. A100 still carried a large\-cluster premium\. The small sample is partly the point: a headline GPU price can rest on a seemingly active market while the market for one specific usable configuration is almost empty\.
If you enjoy this newsletter, consider sharing it with a colleague\.
[Share Buy the Rumor; Sell the News](https://davefriedman.substack.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share)
I’m always happy to receive comments, questions, and pushback\. If you want to connect with me directly, you can:
- **follow me**on[Twitter](https://x.com/friedmandave),
- **connect with me**on[LinkedIn](https://www.linkedin.com/in/frieddave/), or
- **send an email**to dave \[at\] davefriedman dot co\. \(Not \.com\!\)