Home ›Insights ›Kubernetes
KubernetesCost EngineeringOctober 2, 202613 min read

Kubernetes Cost Optimisation: Right-Sizing, Autoscaling and Spot in 2026

Average CPU utilisation across production Kubernetes clusters is 8% — down from 10% the year before. More tooling, more FinOps teams, more blog posts about this exact subject, and the number went backwards. That tells you the problem is not a tooling problem.

RP
Ravi Pratap Singh
Cloud & DevOps Engineering · Vikgol
Share
Kubernetes clusters average 8 percent CPU utilisation and 20 percent memory utilisation in 2026, with 70 percent of requested resources wasted

Cast AI's 2026 State of Kubernetes Optimization Report, measured across more than 23,000 production clusters, found average CPU utilisation at 8% and memory at 20%. Both figures fell from the previous year, when CPU sat at 10%.

That is worth pausing on. The tooling got better. Karpenter matured. Every cloud added cost dashboards. FinOps became a recognised function. And clusters ran less efficiently than they did twelve months earlier.

Which means the standard advice — install a cost tool, look at the dashboard, right-size some pods — is not working. Not because it is wrong, but because it treats the symptom.

📌 Quick answer: why Kubernetes costs so much

The scheduler reserves capacity equal to a pod's requests, not its actual usage. A pod requesting 2 CPU and using 0.1 CPU occupies 2 CPU of node capacity regardless. Multiply across hundreds of pods and you are paying for nodes that are nearly empty by any measure except the one the scheduler uses. Right-sizing requests against observed usage is the single highest-impact fix, typically recovering 30–50% of cluster cost.

The Mechanism Most Teams Never Internalise

Kubernetes has two numbers per container and they do very different things.

YAML — the two numbers that decide your bill
# What the scheduler reserves — this is what you pay for
resources:
  requests:
    cpu: "2000m"      # 2 full cores held for this pod
    memory: "4Gi"

  # The ceiling before throttling or OOM kill
  limits:
    cpu: "4000m"
    memory: "8Gi"

# Observed over 30 days in production:
#   p95 CPU    = 180m   → you reserved 2000m
#   p95 memory = 900Mi  → you reserved 4Gi
#
# You are paying for 2 cores to run a tenth of one.

No alert fires for this. The pod is healthy, the service is up, dashboards are green. The waste is completely invisible from inside the application — it only shows up on the invoice, aggregated with everything else.

Why the Gap Keeps Widening

The incentives explain the data better than any technical account does.

⚠️ Why requests drift upward
Rational individual behaviour
  • · An engineer who over-requests is never paged
  • · An engineer who under-requests causes an incident
  • · Requests get copied from an older, bigger service
  • · A number set during a load test is never revisited
  • · Nobody sees the cost of their own namespace
  • · Lowering a request has no reward and real downside
✅ What actually reverses it
Changed incentives, then tooling
  • · Per-namespace cost visible to the team that owns it
  • · Requests derived from observed p95, not guessed
  • · VPA recommendations in the pull request
  • · A utilisation target that is explicitly not 100%
  • · Someone accountable for cluster efficiency
  • · Review cadence, because drift is continuous

Installing a cost tool without fixing the incentive produces a dashboard nobody acts on. That is the most plausible explanation for utilisation falling in a year when tooling improved.

Where the Money Actually Goes

Cost driverShare of cluster spendPrimary leverCommon waste pattern
Compute — worker nodes60–80%Right-sizing, bin packing, spotOverprovisioned requests, idle nodes held warm by stale reservations
GPU nodesVaries, growing fastGPU right-sizing, time-slicing, spot fallbackLightweight models on dedicated GPUs at near-zero utilisation
Storage — PVs, snapshots10–20%Lifecycle policies, orphan cleanupVolumes from deleted workloads nobody reclaimed
Network — egress, LBs5–15%Topology-aware routing, LB consolidationCross-AZ chatter, one load balancer per service

Compute dominates, which is why right-sizing comes first and everything else is a second-order optimisation.

⚠️ The fastest-growing waste category

GPU idle time. NVIDIA's own documentation notes 0–10% GPU compute utilisation for lightweight models running on dedicated GPUs — which is exactly what happens when a team schedules a small inference workload onto a full A100 or H100 because that is what the node pool offers. GPU time-slicing and MIG partitioning exist precisely for this, and most clusters running AI workloads have neither configured.

Stop Chasing 100% Utilisation

A common overcorrection: a team discovers 8% utilisation, panics, and sets a target of 90%. Then latency degrades, pods get OOM-killed during traffic spikes, and the whole initiative gets reversed with a note saying cost optimisation hurt reliability.

The right target depends entirely on workload shape.

Workload typeTarget node utilisationWhy
Stateless services with HPA70–80%Horizontal scaling absorbs spikes, so headroom can be thin
Bursty or event-driven workloads50–60%Spikes arrive faster than new nodes can be provisioned
Strict latency requirements40–50%CPU throttling at high utilisation shows up directly in p99
Batch and async processing80–90%Delay is acceptable; this is also the best spot candidate

Getting from 8% to 55% is a transformational result. Getting from 8% to 90% is an outage waiting for a traffic spike.

The Four Levers, in Order

LEVER 01 — DO THIS FIRST
Right-size requests against p95
Set requests to observed 95th-percentile usage plus 20–30% headroom, not to whatever was guessed at deploy time. Typically recovers 30–50% of cluster cost and takes one to two weeks for a mid-size cluster. Nothing else comes close on return per hour spent.
LEVER 02
Bin-pack onto fewer, better-matched nodes
Once requests are honest, the same workloads fit on fewer nodes. Multiple node groups — general, memory-optimised, spot — let the scheduler pick the right shape instead of forcing everything onto general-purpose instances.
LEVER 03
Move fault-tolerant work to spot
Spot is 60–90% cheaper for equivalent compute. Batch jobs, async workers, CI runners, stateless replicas behind a load balancer. Requires Pod Disruption Budgets and multiple capacity pools — without those it is a reliability problem, not a saving.
LEVER 04
Commit the stable baseline
Whatever capacity you demonstrably run every hour of every day should be on reserved instances or a savings plan. This is the last lever, not the first — commit before right-sizing and you lock in your own waste for three years.
✅ Why the order matters more than the levers

Each lever multiplies the ones before it. Right-sizing first means you bin-pack honest numbers, buy spot capacity you actually need, and commit to a baseline that reflects real consumption. Doing it in reverse — committing first, then right-sizing — means you have already paid for three years of the waste you were about to eliminate. We have seen teams do exactly this, and the reserved instance contract is the one part that cannot be undone.

Spot: What Actually Works

Spot instances are the largest single discount available and the most commonly misapplied. The failure mode is treating spot as a cheaper on-demand rather than a different reliability contract.

✅ Safe on spot
Interruption-tolerant
  • · Batch and ETL jobs with checkpointing
  • · CI/CD runners
  • · Async queue workers
  • · Stateless replicas with a sensible PDB
  • · Dev and staging environments entirely
  • · ML training with checkpoint and resume
⚠️ Not safe on spot
Stateful or latency-critical
  • · Databases and stateful sets
  • · Message brokers holding unreplicated state
  • · Services with sticky sessions
  • · Anything where a 2-minute eviction breaks an SLA
  • · Single-replica control plane components
  • · Long-running jobs with no checkpointing

Three things make spot work in practice: Pod Disruption Budgets so evictions are orderly, multiple instance types and availability zones so one capacity pool drying up does not take the workload with it, and on-demand fallback so the cluster degrades to paying more rather than to not running.

Autoscaling: The Interaction Nobody Tests

Three autoscalers, operating on different signals, frequently working against each other:

  • HPA adds pod replicas based on observed metrics
  • VPA adjusts requests and limits on existing pods
  • Cluster Autoscaler or Karpenter adds and removes nodes based on pending pods

Running HPA and VPA on the same CPU metric is a known conflict — VPA raises the request, which lowers observed utilisation percentage, which tells HPA to scale down, which raises per-pod load, which prompts VPA again. Use VPA in recommendation mode alongside HPA rather than both acting on the same signal.

Karpenter changes the node-side calculation meaningfully: it provisions instances matched to pending pod shapes and continuously consolidates underutilised nodes, rather than scaling predefined node groups. It typically adds 10–20% on top of right-sizing — but it will over-provision without explicit constraints, and it will evict pods that have no Pod Disruption Budget.

A 30-60-90 Plan

1

Measure the request-to-usage gap

Per namespace, per workload: requested CPU and memory against observed p95 over 30 days. Express it as a waste percentage. Above 60% means no right-sizing has ever happened; 20–40% is decent hygiene. This single table drives everything after it.

Days 1–10
2

Make cost visible per namespace

OpenCost or Kubecost, broken down by team. Not to charge anyone back on day one — just so the team that owns a namespace can see what it costs. Visibility alone moves behaviour before any policy does.

Days 1–15
3

Right-size the top 20 workloads

Sorted by absolute waste, not by percentage. A 90% waste rate on a tiny service is irrelevant; a 60% waste rate on your largest deployment is the whole project. Deploy to staging, watch for throttling and OOM kills, then promote.

Days 10–30
4

Introduce spot for a non-production workload first

Dev and staging entirely on spot, with PDBs and multiple instance types configured. This builds the operational muscle before anything production-facing depends on it. Then move CI runners, then async workers.

Days 30–60
5

Add node consolidation

Karpenter or equivalent, with explicit instance-type constraints and disruption budgets. Expect 10–20% beyond right-sizing. Do not deploy it to a cluster that has not been right-sized — it will faithfully consolidate your overprovisioned requests.

Days 45–75
6

Commit the baseline, then institutionalise review

Only now buy reserved capacity for what you demonstrably run continuously. Then set a monthly review of the waste table, because requests drift upward by default — this is maintenance, not a project with an end date.

Days 75–90

Frequently Asked Questions

How much can Kubernetes cost optimisation realistically save?
Teams that have never done it typically see 40–60% reduction from right-sizing plus spot together. Cast AI's benchmark across production clusters shows around 43% average compute cost reduction when combining right-sizing, spot and automated node management. Right-sizing alone usually accounts for 30–50%. Teams with mature practices can reach 70% below a naive baseline, but the first two levers deliver most of it — the long tail is genuinely diminishing returns.
Why is cluster utilisation so low in the first place?
Because the scheduler reserves what you request, not what you use, and the incentives all push requests upward. An engineer who over-requests is never paged; one who under-requests causes an incident. Requests get copied from older services, set during a load test and never revisited, and nobody sees the cost of their own namespace. Lowering a request carries real downside and no reward. That is why utilisation fell from 10% to 8% in a year when tooling improved — the tooling was never the binding constraint.
What utilisation should we actually target?
It depends on workload shape, and chasing a single number across a cluster is how cost optimisation gets blamed for an outage. Stateless services with HPA: 70–80%. Bursty or event-driven: 50–60%, because spikes arrive faster than nodes can be provisioned. Strict latency requirements: 40–50%, because CPU throttling at high utilisation lands directly in p99. Batch and async: 80–90%. Moving from 8% to 55% is already a transformational result.
Which workloads are safe on spot instances?
Anything that tolerates a two-minute eviction notice: batch jobs with checkpointing, CI runners, async queue workers, stateless replicas behind a load balancer, and entire dev and staging environments. Not safe: databases, stateful sets, message brokers holding unreplicated state, services with sticky sessions, and anything where eviction breaks an SLA. Three things make it work — Pod Disruption Budgets, multiple instance types and availability zones, and on-demand fallback.
Should we use VPA and HPA together?
Not on the same metric. VPA raising a CPU request lowers the observed utilisation percentage, which tells HPA to scale down, which raises per-pod load, which prompts VPA again. The standard pattern is VPA in recommendation mode — surfacing suggested requests to engineers rather than applying them — alongside HPA acting on actual scaling decisions. Or HPA on a custom metric like queue depth, with VPA managing resources independently.
What is the biggest Kubernetes cost mistake in 2026?
Running AI workloads on on-demand GPUs with no spot fallback and no time-slicing. GPU idle time is the fastest-growing waste category — NVIDIA documents 0–10% compute utilisation for lightweight models on dedicated GPUs, which is exactly what happens when a small inference workload lands on a full H100 because that is what the node pool offers. GPU time-slicing and MIG partitioning exist for this, and most clusters running AI have neither configured.

Want to Know Your Actual Waste Number?

We do AWS architecture, Kubernetes cost optimisation and DR/BCP for production systems. Book a free 30-minute call — we'll look at your request-to-usage gap and tell you honestly what is recoverable and what is not.

#Kubernetes#K8s#FinOps#CloudCost#Karpenter#SpotInstances#DevOps#Vikgol
RP
Ravi Pratap Singh
Cloud & DevOps Engineering · Vikgol
Ravi works on production Kubernetes and AWS infrastructure at Vikgol, including cost optimisation, disaster recovery and business continuity for regulated financial services clients. Vikgol has shipped 90+ AI, web and cloud projects for startups and enterprises across US, UK, UAE and India.
Available Now · 72-Hour POC

Ready to Ship Your AI Product?
Let’s Build It Together.

Senior engineers on demand. Working prototype in 72 hours. NDA before we discuss anything. 100% code ownership to you — no lock-in, ever.

70+
Senior engineers on staff
90+
Projects delivered globally
72h
Working POC guaranteed
5★
Client satisfaction rating