In July 2026, an NVIDIA H100 80GB rented for as little as $1.38 per GPU-hour on marketplace and boutique clouds, and as much as $12.29 on Azure's on-demand instances. Google Cloud listed its eight-GPU H100 node at $88.49 an hour, roughly $11.06 per GPU.
Same chip. Same memory. Twelvefold spread.
That gap alone is enough to make or break an AI product's unit economics. But it is not the most expensive mistake teams make, and a blog about GPU pricing that stops at comparing hourly rates has skipped the part that matters.
The most expensive GPU decision is usually renting one you did not need.
Three numbers, in order of impact. Utilisation — what fraction of the hours you pay for are doing useful work. Pricing model — per-token API, serverless, on-demand, spot, or reserved. Hourly rate — which provider and which GPU. Most cost conversations start at the third and never reach the first, which is backwards.
The Price Spread Is Real, and Exploitable
| Where you rent it | H100 per GPU-hour | What you give up |
|---|---|---|
| Spot / interruptible | $1.25 – $1.49 | Reclaimed with 30 seconds to 2 minutes notice |
| Marketplace (Vast.ai, similar) | $1.38 – $2.00 | Host variability, no SLA, uneven reliability |
| Neo-clouds (Lambda, Runpod, Spheron) | $2.00 – $3.00 | Smaller regional footprint, fewer managed services |
| Market average, early 2026 | ~$3.11 | — |
| AWS P5 on-demand | ~$12.29 | Nothing — you are paying for ecosystem and SLA |
| Azure ND H100 v5 | $11 – $13 | Nothing — same |
The hyperscaler premium is not irrational. You get regional coverage, compliance certifications, integration with the rest of your stack, and an SLA someone will honour. For a regulated fintech running inference on customer data inside an existing AWS estate, that is worth paying for.
For a batch training job that can checkpoint and resume, it is not. The mistake is applying one answer to both.
Data egress on hyperscalers runs $0.08–$0.12 per GB. Moving a 2TB training dataset out of AWS costs over $160 before you have done any compute. For multi-cloud setups — training on one provider, serving on another — egress can quietly exceed the GPU bill. Model it before you architect around it.
Utilisation Beats Hourly Rate, Consistently
Here is the arithmetic that should govern the decision, and rarely does.
An H100 at $3.99 an hour, running continuously, costs roughly $35,000 per GPU-year — more than buying the card outright. An eight-GPU node runs $49 an hour at CoreWeave or $55 at AWS: $36,000 to $40,000 a month at full utilisation.
Now apply the number nobody wants to measure. Most teams running dedicated GPU instances are somewhere between 15% and 40% utilised. At 25% utilisation, you are paying four times the effective rate — which means a $1.38 GPU sitting idle three-quarters of the time costs more per unit of work than a $3.00 GPU running flat out.
Below roughly 30% utilisation, serverless or per-request billing wins because you stop paying for idle. Above it, dedicated hourly billing wins because you avoid cold starts. That single threshold answers most of the architecture question — and you cannot evaluate it without measuring your actual utilisation first.
The Cheapest GPU Is Often No GPU
This is the part GPU providers do not write blog posts about.
A large share of inference workloads running on rented H100s would cost a fraction as much on a per-token API. The gap widened sharply in 2026: on 30 July, OpenAI cut GPT-5.6 Luna pricing by 80% and Terra by 20%, while GPU rental rates fell far less.
If you are running a standard open-weight model at moderate volume, doing the comparison honestly is often uncomfortable. The maths frequently favours the API by a wide margin, and the team has already provisioned the cluster.
- · Traffic is variable or unpredictable
- · You are using a standard model, not a custom one
- · Volume is moderate — below a few million tokens a day
- · You have no MLOps team to run the stack
- · Time to market matters more than cost per token
- · Your data can leave your environment
- · Sustained high volume with predictable load
- · Fine-tuned or proprietary model weights
- · Data cannot leave your infrastructure — regulated sectors
- · Latency requirements an API cannot meet
- · You are training, not just serving
- · You have measured utilisation above 60%
Take last month's actual token volume. Price it on a per-token API at current rates. Compare against your GPU bill including idle hours, storage and egress. If the API is cheaper — and for many teams it is — the GPU cluster is buying you something other than cost efficiency. That something might be genuinely necessary, like data residency. Be able to name it.
Which GPU, If You Do Need One
| GPU | Best for | Rough on-demand range |
|---|---|---|
| A100 80GB | Smaller and quantised models, cost-sensitive inference | $1.00 – $2.70 |
| H100 80GB | General-purpose training and 7B–70B inference. Mature toolchain, best price-performance at 2026 rates | $1.38 – $12.29 |
| H200 141GB | Memory-bound inference — long context, large KV cache. Premium usually pays here | Between H100 and B200 |
| B200 | Frontier training and 100B+ inference with heavy concurrency | $2.12 spot – $14.24 on AWS |
| L4 / T4 | Light inference, embedding generation, non-LLM workloads | Under $1.00 |
The decision rule compresses to one sentence: memory-bound inference favours H200, training throughput favours B200 per result, and everything cost-sensitive with a mature toolchain favours H100 at current rates.
Worth knowing about the market itself: Blackwell shipping has pushed Hopper-generation GPUs into a price decline. H100 spot prices in some AWS regions fell as much as 88% between January 2024 and September 2025. H200 spot inventory stays thin, so it commands a premium. Blackwell availability remains constrained through 2026.
Buy Versus Rent
The crossover arrives inside a year for genuinely sustained workloads. An H100 at $3.99/hour is about $35,000 per GPU-year at full utilisation, against roughly $25,000–$40,000 to buy the card — before the server, networking, power, cooling and someone to operate it.
A 30-Day Cost Review
Measure actual utilisation
GPU-hours billed against GPU-hours doing useful work, per workload. This single number changes more decisions than any pricing comparison. Most teams are surprised, and not pleasantly.
Week 1Price your volume on a per-token API
Take last month's real token counts and price them at current API rates. Compare honestly against the GPU bill including idle time, storage and egress. Do this before optimising anything else.
Week 1Split workloads by interruption tolerance
Batch and training work moves to spot at 40–65% savings. Production endpoints stay on-demand or reserved. Many teams run everything on one tier because it was simpler to set up, and pay for that simplicity monthly.
Week 2Fix batching and quantisation before adding GPUs
Continuous batching and FP8 quantisation routinely double or triple throughput per GPU. Capacity problems are frequently efficiency problems wearing a disguise, and adding hardware makes the disguise permanent.
Week 2–3Model egress before going multi-cloud
Training on one provider and serving on another sounds clean architecturally. At $0.08–$0.12 per GB it can quietly cost more than the compute. Private cross-cloud networking cuts this 60–80% but needs explicit setup.
Week 3Instrument cost per unit of business value
Cost per thousand tokens, per document processed, per conversation handled. A GPU bill in isolation tells you nothing about whether it is too high. Cost per outcome tells you immediately.
OngoingFrequently Asked Questions
Paying for GPUs You're Not Using?
We do AWS architecture, AI infrastructure and cost optimisation. Book a free 30-minute call — we'll look at your utilisation and tell you honestly whether the fix is cheaper GPUs, better batching, or no GPUs at all.