Cast AI's 2026 State of Kubernetes Optimization Report, measured across more than 23,000 production clusters, found average CPU utilisation at 8% and memory at 20%. Both figures fell from the previous year, when CPU sat at 10%.
That is worth pausing on. The tooling got better. Karpenter matured. Every cloud added cost dashboards. FinOps became a recognised function. And clusters ran less efficiently than they did twelve months earlier.
Which means the standard advice — install a cost tool, look at the dashboard, right-size some pods — is not working. Not because it is wrong, but because it treats the symptom.
The scheduler reserves capacity equal to a pod's requests, not its actual usage. A pod requesting 2 CPU and using 0.1 CPU occupies 2 CPU of node capacity regardless. Multiply across hundreds of pods and you are paying for nodes that are nearly empty by any measure except the one the scheduler uses. Right-sizing requests against observed usage is the single highest-impact fix, typically recovering 30–50% of cluster cost.
The Mechanism Most Teams Never Internalise
Kubernetes has two numbers per container and they do very different things.
# What the scheduler reserves — this is what you pay for resources: requests: cpu: "2000m" # 2 full cores held for this pod memory: "4Gi" # The ceiling before throttling or OOM kill limits: cpu: "4000m" memory: "8Gi" # Observed over 30 days in production: # p95 CPU = 180m → you reserved 2000m # p95 memory = 900Mi → you reserved 4Gi # # You are paying for 2 cores to run a tenth of one.
No alert fires for this. The pod is healthy, the service is up, dashboards are green. The waste is completely invisible from inside the application — it only shows up on the invoice, aggregated with everything else.
Why the Gap Keeps Widening
The incentives explain the data better than any technical account does.
- · An engineer who over-requests is never paged
- · An engineer who under-requests causes an incident
- · Requests get copied from an older, bigger service
- · A number set during a load test is never revisited
- · Nobody sees the cost of their own namespace
- · Lowering a request has no reward and real downside
- · Per-namespace cost visible to the team that owns it
- · Requests derived from observed p95, not guessed
- · VPA recommendations in the pull request
- · A utilisation target that is explicitly not 100%
- · Someone accountable for cluster efficiency
- · Review cadence, because drift is continuous
Installing a cost tool without fixing the incentive produces a dashboard nobody acts on. That is the most plausible explanation for utilisation falling in a year when tooling improved.
Where the Money Actually Goes
| Cost driver | Share of cluster spend | Primary lever | Common waste pattern |
|---|---|---|---|
| Compute — worker nodes | 60–80% | Right-sizing, bin packing, spot | Overprovisioned requests, idle nodes held warm by stale reservations |
| GPU nodes | Varies, growing fast | GPU right-sizing, time-slicing, spot fallback | Lightweight models on dedicated GPUs at near-zero utilisation |
| Storage — PVs, snapshots | 10–20% | Lifecycle policies, orphan cleanup | Volumes from deleted workloads nobody reclaimed |
| Network — egress, LBs | 5–15% | Topology-aware routing, LB consolidation | Cross-AZ chatter, one load balancer per service |
Compute dominates, which is why right-sizing comes first and everything else is a second-order optimisation.
GPU idle time. NVIDIA's own documentation notes 0–10% GPU compute utilisation for lightweight models running on dedicated GPUs — which is exactly what happens when a team schedules a small inference workload onto a full A100 or H100 because that is what the node pool offers. GPU time-slicing and MIG partitioning exist precisely for this, and most clusters running AI workloads have neither configured.
Stop Chasing 100% Utilisation
A common overcorrection: a team discovers 8% utilisation, panics, and sets a target of 90%. Then latency degrades, pods get OOM-killed during traffic spikes, and the whole initiative gets reversed with a note saying cost optimisation hurt reliability.
The right target depends entirely on workload shape.
| Workload type | Target node utilisation | Why |
|---|---|---|
| Stateless services with HPA | 70–80% | Horizontal scaling absorbs spikes, so headroom can be thin |
| Bursty or event-driven workloads | 50–60% | Spikes arrive faster than new nodes can be provisioned |
| Strict latency requirements | 40–50% | CPU throttling at high utilisation shows up directly in p99 |
| Batch and async processing | 80–90% | Delay is acceptable; this is also the best spot candidate |
Getting from 8% to 55% is a transformational result. Getting from 8% to 90% is an outage waiting for a traffic spike.
The Four Levers, in Order
Each lever multiplies the ones before it. Right-sizing first means you bin-pack honest numbers, buy spot capacity you actually need, and commit to a baseline that reflects real consumption. Doing it in reverse — committing first, then right-sizing — means you have already paid for three years of the waste you were about to eliminate. We have seen teams do exactly this, and the reserved instance contract is the one part that cannot be undone.
Spot: What Actually Works
Spot instances are the largest single discount available and the most commonly misapplied. The failure mode is treating spot as a cheaper on-demand rather than a different reliability contract.
- · Batch and ETL jobs with checkpointing
- · CI/CD runners
- · Async queue workers
- · Stateless replicas with a sensible PDB
- · Dev and staging environments entirely
- · ML training with checkpoint and resume
- · Databases and stateful sets
- · Message brokers holding unreplicated state
- · Services with sticky sessions
- · Anything where a 2-minute eviction breaks an SLA
- · Single-replica control plane components
- · Long-running jobs with no checkpointing
Three things make spot work in practice: Pod Disruption Budgets so evictions are orderly, multiple instance types and availability zones so one capacity pool drying up does not take the workload with it, and on-demand fallback so the cluster degrades to paying more rather than to not running.
Autoscaling: The Interaction Nobody Tests
Three autoscalers, operating on different signals, frequently working against each other:
- HPA adds pod replicas based on observed metrics
- VPA adjusts requests and limits on existing pods
- Cluster Autoscaler or Karpenter adds and removes nodes based on pending pods
Running HPA and VPA on the same CPU metric is a known conflict — VPA raises the request, which lowers observed utilisation percentage, which tells HPA to scale down, which raises per-pod load, which prompts VPA again. Use VPA in recommendation mode alongside HPA rather than both acting on the same signal.
Karpenter changes the node-side calculation meaningfully: it provisions instances matched to pending pod shapes and continuously consolidates underutilised nodes, rather than scaling predefined node groups. It typically adds 10–20% on top of right-sizing — but it will over-provision without explicit constraints, and it will evict pods that have no Pod Disruption Budget.
A 30-60-90 Plan
Measure the request-to-usage gap
Per namespace, per workload: requested CPU and memory against observed p95 over 30 days. Express it as a waste percentage. Above 60% means no right-sizing has ever happened; 20–40% is decent hygiene. This single table drives everything after it.
Days 1–10Make cost visible per namespace
OpenCost or Kubecost, broken down by team. Not to charge anyone back on day one — just so the team that owns a namespace can see what it costs. Visibility alone moves behaviour before any policy does.
Days 1–15Right-size the top 20 workloads
Sorted by absolute waste, not by percentage. A 90% waste rate on a tiny service is irrelevant; a 60% waste rate on your largest deployment is the whole project. Deploy to staging, watch for throttling and OOM kills, then promote.
Days 10–30Introduce spot for a non-production workload first
Dev and staging entirely on spot, with PDBs and multiple instance types configured. This builds the operational muscle before anything production-facing depends on it. Then move CI runners, then async workers.
Days 30–60Add node consolidation
Karpenter or equivalent, with explicit instance-type constraints and disruption budgets. Expect 10–20% beyond right-sizing. Do not deploy it to a cluster that has not been right-sized — it will faithfully consolidate your overprovisioned requests.
Days 45–75Commit the baseline, then institutionalise review
Only now buy reserved capacity for what you demonstrably run continuously. Then set a monthly review of the waste table, because requests drift upward by default — this is maintenance, not a project with an end date.
Days 75–90Frequently Asked Questions
Want to Know Your Actual Waste Number?
We do AWS architecture, Kubernetes cost optimisation and DR/BCP for production systems. Book a free 30-minute call — we'll look at your request-to-usage gap and tell you honestly what is recoverable and what is not.