Home ›Insights ›AI Infrastructure
AI InfrastructureCost EngineeringSeptember 26, 202613 min read

Running AI Workloads in the Cloud: GPU Cost and Scaling in 2026

The same NVIDIA H100 rents for around $1.38 per GPU-hour on a marketplace and $12.29 on Azure. Identical silicon, identical memory, a twelvefold spread — the only variable is who is selling it. And that spread is still not the most expensive decision most teams get wrong.

VE
Vikgol Engineering Team
Cloud & AI Infrastructure · Vikgol
Share
The same H100 GPU rents from $1.38 to $12.29 per hour depending on provider, and utilisation matters more than the hourly rate

In July 2026, an NVIDIA H100 80GB rented for as little as $1.38 per GPU-hour on marketplace and boutique clouds, and as much as $12.29 on Azure's on-demand instances. Google Cloud listed its eight-GPU H100 node at $88.49 an hour, roughly $11.06 per GPU.

Same chip. Same memory. Twelvefold spread.

That gap alone is enough to make or break an AI product's unit economics. But it is not the most expensive mistake teams make, and a blog about GPU pricing that stops at comparing hourly rates has skipped the part that matters.

The most expensive GPU decision is usually renting one you did not need.

📌 Quick answer: what drives GPU cost?

Three numbers, in order of impact. Utilisation — what fraction of the hours you pay for are doing useful work. Pricing model — per-token API, serverless, on-demand, spot, or reserved. Hourly rate — which provider and which GPU. Most cost conversations start at the third and never reach the first, which is backwards.

The Price Spread Is Real, and Exploitable

Where you rent itH100 per GPU-hourWhat you give up
Spot / interruptible$1.25 – $1.49Reclaimed with 30 seconds to 2 minutes notice
Marketplace (Vast.ai, similar)$1.38 – $2.00Host variability, no SLA, uneven reliability
Neo-clouds (Lambda, Runpod, Spheron)$2.00 – $3.00Smaller regional footprint, fewer managed services
Market average, early 2026~$3.11—
AWS P5 on-demand~$12.29Nothing — you are paying for ecosystem and SLA
Azure ND H100 v5$11 – $13Nothing — same

The hyperscaler premium is not irrational. You get regional coverage, compliance certifications, integration with the rest of your stack, and an SLA someone will honour. For a regulated fintech running inference on customer data inside an existing AWS estate, that is worth paying for.

For a batch training job that can checkpoint and resume, it is not. The mistake is applying one answer to both.

⚠️ The cost that does not appear in any hourly rate

Data egress on hyperscalers runs $0.08–$0.12 per GB. Moving a 2TB training dataset out of AWS costs over $160 before you have done any compute. For multi-cloud setups — training on one provider, serving on another — egress can quietly exceed the GPU bill. Model it before you architect around it.

Utilisation Beats Hourly Rate, Consistently

Here is the arithmetic that should govern the decision, and rarely does.

An H100 at $3.99 an hour, running continuously, costs roughly $35,000 per GPU-year — more than buying the card outright. An eight-GPU node runs $49 an hour at CoreWeave or $55 at AWS: $36,000 to $40,000 a month at full utilisation.

Now apply the number nobody wants to measure. Most teams running dedicated GPU instances are somewhere between 15% and 40% utilised. At 25% utilisation, you are paying four times the effective rate — which means a $1.38 GPU sitting idle three-quarters of the time costs more per unit of work than a $3.00 GPU running flat out.

~30%
Utilisation break-even between serverless and dedicated hourly billing
Spheron, 2026
$35K
Annual cost of one rented H100 at $3.99/hr, 24/7 — more than buying the card
CloudZero
40–65%
Typical spot savings against on-demand, where interruption is tolerable
Multiple providers

Below roughly 30% utilisation, serverless or per-request billing wins because you stop paying for idle. Above it, dedicated hourly billing wins because you avoid cold starts. That single threshold answers most of the architecture question — and you cannot evaluate it without measuring your actual utilisation first.

The Cheapest GPU Is Often No GPU

This is the part GPU providers do not write blog posts about.

A large share of inference workloads running on rented H100s would cost a fraction as much on a per-token API. The gap widened sharply in 2026: on 30 July, OpenAI cut GPT-5.6 Luna pricing by 80% and Terra by 20%, while GPU rental rates fell far less.

If you are running a standard open-weight model at moderate volume, doing the comparison honestly is often uncomfortable. The maths frequently favours the API by a wide margin, and the team has already provisioned the cluster.

⚡ Use a per-token API when
Most teams, most of the time
  • · Traffic is variable or unpredictable
  • · You are using a standard model, not a custom one
  • · Volume is moderate — below a few million tokens a day
  • · You have no MLOps team to run the stack
  • · Time to market matters more than cost per token
  • · Your data can leave your environment
🖥️ Rent or buy GPUs when
Specific, defensible reasons
  • · Sustained high volume with predictable load
  • · Fine-tuned or proprietary model weights
  • · Data cannot leave your infrastructure — regulated sectors
  • · Latency requirements an API cannot meet
  • · You are training, not just serving
  • · You have measured utilisation above 60%
✅ The honest test

Take last month's actual token volume. Price it on a per-token API at current rates. Compare against your GPU bill including idle hours, storage and egress. If the API is cheaper — and for many teams it is — the GPU cluster is buying you something other than cost efficiency. That something might be genuinely necessary, like data residency. Be able to name it.

Which GPU, If You Do Need One

GPUBest forRough on-demand range
A100 80GBSmaller and quantised models, cost-sensitive inference$1.00 – $2.70
H100 80GBGeneral-purpose training and 7B–70B inference. Mature toolchain, best price-performance at 2026 rates$1.38 – $12.29
H200 141GBMemory-bound inference — long context, large KV cache. Premium usually pays hereBetween H100 and B200
B200Frontier training and 100B+ inference with heavy concurrency$2.12 spot – $14.24 on AWS
L4 / T4Light inference, embedding generation, non-LLM workloadsUnder $1.00

The decision rule compresses to one sentence: memory-bound inference favours H200, training throughput favours B200 per result, and everything cost-sensitive with a mature toolchain favours H100 at current rates.

Worth knowing about the market itself: Blackwell shipping has pushed Hopper-generation GPUs into a price decline. H100 spot prices in some AWS regions fell as much as 88% between January 2024 and September 2025. H200 spot inventory stays thin, so it commands a premium. Blackwell availability remains constrained through 2026.

Buy Versus Rent

The crossover arrives inside a year for genuinely sustained workloads. An H100 at $3.99/hour is about $35,000 per GPU-year at full utilisation, against roughly $25,000–$40,000 to buy the card — before the server, networking, power, cooling and someone to operate it.

RENT WHEN
Utilisation is uncertain
Which is most of the time, especially in the first year of a product. Renting buys optionality, and optionality is worth a premium when your load forecast is really a guess.
BUY WHEN
You have measured demand, not forecast it
Pull three months of actual GPU-hours from rental invoices. Not a capacity plan from a roadmap deck. Teams that regret buying did not get the hardware maths wrong — they believed the optimistic version of their own utilisation forecast.
RESERVE WHEN
Load is steady but you do not want hardware
36-month reserved contracts reach as low as $2.25 per GPU-hour for B200-class capacity. Most of the buying discount, none of the operational burden — but you are committed for three years in a market that moves every quarter.
USE SPOT WHEN
The work can be interrupted
Batch training with checkpoint and resume, offline inference pipelines, hyperparameter sweeps, data preprocessing. 40–65% savings. Never for a production endpoint with an availability commitment.

A 30-Day Cost Review

1

Measure actual utilisation

GPU-hours billed against GPU-hours doing useful work, per workload. This single number changes more decisions than any pricing comparison. Most teams are surprised, and not pleasantly.

Week 1
2

Price your volume on a per-token API

Take last month's real token counts and price them at current API rates. Compare honestly against the GPU bill including idle time, storage and egress. Do this before optimising anything else.

Week 1
3

Split workloads by interruption tolerance

Batch and training work moves to spot at 40–65% savings. Production endpoints stay on-demand or reserved. Many teams run everything on one tier because it was simpler to set up, and pay for that simplicity monthly.

Week 2
4

Fix batching and quantisation before adding GPUs

Continuous batching and FP8 quantisation routinely double or triple throughput per GPU. Capacity problems are frequently efficiency problems wearing a disguise, and adding hardware makes the disguise permanent.

Week 2–3
5

Model egress before going multi-cloud

Training on one provider and serving on another sounds clean architecturally. At $0.08–$0.12 per GB it can quietly cost more than the compute. Private cross-cloud networking cuts this 60–80% but needs explicit setup.

Week 3
6

Instrument cost per unit of business value

Cost per thousand tokens, per document processed, per conversation handled. A GPU bill in isolation tells you nothing about whether it is too high. Cost per outcome tells you immediately.

Ongoing

Frequently Asked Questions

Why does the same H100 cost $1.38 on one provider and $12.29 on another?
You are buying different things. The cheap end is marketplace and interruptible capacity — no SLA, variable host quality, reclaimable at short notice. The expensive end is hyperscaler on-demand with regional coverage, compliance certifications, integration with the rest of your stack, and a guarantee someone will honour. Neither is wrong. Applying one answer to every workload is.
Should we use a per-token API or run our own GPUs?
Start with the API unless you have a specific reason not to. Run your own GPUs when you have sustained high volume with predictable load, fine-tuned or proprietary weights, data that cannot leave your infrastructure, latency an API cannot meet, or you are training rather than serving. If none of those apply, the API is almost certainly cheaper — and after OpenAI's July 2026 price cuts of 80% on Luna and 20% on Terra, the gap widened further while GPU rates barely moved.
What GPU utilisation should we be targeting?
Above 60% before dedicated instances make economic sense against serverless. The break-even sits around 30% — below that, per-second serverless billing wins because you stop paying for idle; above it, dedicated wins because you avoid cold starts. Most teams running dedicated GPUs are between 15% and 40% and have never measured it. That measurement is the highest-return hour in any GPU cost review.
Is it worth buying GPUs instead of renting?
Only with measured, sustained demand. At $3.99/hour a rented H100 costs about $35,000 per GPU-year at full utilisation, against roughly $25,000–$40,000 to buy — so the crossover arrives inside a year in theory. In practice you also need the server, networking, power, cooling and an operator. The teams that regret buying did not get the hardware maths wrong; they believed an optimistic utilisation forecast. Pull three months of actual GPU-hours from your invoices — that is the only honest forecast you own.
When does Blackwell make sense over H100?
For frontier model training and 100B-plus inference with heavy concurrency, B200 and GB200 change the throughput-per-dollar maths meaningfully. For most 7B–70B inference today, H100 and H200 still win on price — particularly as Hopper enters a price decline cycle while Blackwell supply stays constrained. If you are memory-bound rather than compute-bound, H200's 141GB is usually the better upgrade than jumping to Blackwell.
What costs do people forget to budget for?
Four, consistently. Data egress at $0.08–$0.12 per GB, which can exceed compute in multi-cloud setups. Storage for datasets and checkpoints, billed separately from GPU-hours. Idle time between jobs, which is the largest and least visible. And engineering time to run the stack — provisioning, monitoring, debugging OOM errors at 2am. That last one is the real argument for managed inference when your team is small.

Paying for GPUs You're Not Using?

We do AWS architecture, AI infrastructure and cost optimisation. Book a free 30-minute call — we'll look at your utilisation and tell you honestly whether the fix is cheaper GPUs, better batching, or no GPUs at all.

#GPUCost#AIInfrastructure#H100#LLMInference#CloudCost#FinOps#AWS#Vikgol
VE
Vikgol Engineering Team
Cloud & AI Infrastructure · Vikgol
The Vikgol engineering team has shipped 90+ AI, web and cloud projects for startups and enterprises across US, UK, UAE and India. We run production AI infrastructure on AWS and are an OpenAI Select Partner — which means we will tell you when the API is cheaper than the cluster.
Available Now · 72-Hour POC

Ready to Ship Your AI Product?
Let’s Build It Together.

Senior engineers on demand. Working prototype in 72 hours. NDA before we discuss anything. 100% code ownership to you — no lock-in, ever.

70+
Senior engineers on staff
90+
Projects delivered globally
72h
Working POC guaranteed
5★
Client satisfaction rating