Cloud

Kubernetes cost optimization: where the money actually goes

Why Kubernetes clusters cost more than they should, how to attribute spend to teams, and the specific changes that produce the largest savings.

Harshit SengarFounder and platform lead, DevOpsArkPublished 31 March 2026 · Updated 3 August 20268 min read
TL;DR

Kubernetes overspend is almost always the gap between what workloads request and what they use, plus idle node capacity nobody owns. Fixing it needs attribution first, because you cannot optimise what you cannot assign, then concrete right-sizing values rather than percentage advice, then cleanup of the orphaned resources that accumulate quietly.

Short answer

Why is Kubernetes so expensive?

Kubernetes costs are usually driven by the gap between resource requests and actual usage. Because the scheduler reserves capacity based on requests, over-provisioned workloads consume cluster capacity they never use, forcing more nodes than the real workload requires. Idle node capacity, orphaned volumes and oversized node pools account for most of the rest.

Requests are the bill, not usage

The central fact of Kubernetes cost is that the scheduler allocates on requests, not on usage. A pod requesting 2 CPU and using 0.1 occupies 2 CPU of schedulable capacity. Ten of those and you are paying for a node to do almost nothing.

This is why cluster utilisation graphs mislead so consistently. A cluster showing 25% CPU usage and refusing to schedule new pods is not a paradox: it is fully allocated and barely used.

Where the numbers usually land. In most estates we see, the aggregate gap between requested and used memory is between two and four times. The requests were copied from another service, or set once during a load test that never recurred.

Attribution comes before optimisation

A cluster has one invoice line and many tenants. Until node cost is distributed across namespaces and workloads, nobody owns any part of it, and shared costs that nobody owns do not get reduced.

Allocation works by distributing each node cost across the pods scheduled onto it, weighted by requests and usage. The important detail is what to do with the unallocated remainder, the capacity that is paid for and not requested by anything. Hiding it in a shared bucket removes the incentive to reduce it; distributing it proportionally makes idle capacity visible to the teams whose over-requesting caused the node to exist.

The changes that produce real savings

ChangeTypical impactEffort
Right-size requests against observed usageLarge, usually the biggest single itemLow per workload, needs data
Remove orphaned volumes, load balancers and IPsModerate and recurringLow
Consolidate under-utilised node poolsModerateMedium, needs a scheduling review
Use spot or preemptible capacity for tolerant workloadsLarge where applicableMedium, needs disruption handling
Scale non-production to zero outside working hoursModerateLow
Commitment coverage for steady-state usageModerateLow, but needs confidence in the baseline
Reduce over-replication of low-traffic servicesSmall individually, adds upLow

Right-sizing without causing an incident

The failure mode of enthusiastic right-sizing is setting limits from average usage and then having pods OOMKilled during a peak. Two rules avoid it.

  • Set requests near the sustained usage percentile you actually observe, commonly p90 over several weeks, including the busiest day of the week.
  • Set memory limits with genuine headroom above peak, because exceeding a memory limit is a hard kill with no grace period. CPU limits are more forgiving, since throttling degrades rather than terminates.
# Observed over 30 days: memory p50 340Mi, p99 620Mi, peak 710Mi.
# CPU p50 0.08, p99 0.31.
resources:
  requests:
    memory: "700Mi"     # near sustained peak, this is what gets scheduled
    cpu: "100m"         # near p50; bursting is fine for CPU
  limits:
    memory: "1Gi"       # headroom above observed peak; exceeding this kills
    # No CPU limit: throttling a latency-sensitive service to save nothing
    # is a common and avoidable self-inflicted problem.

Whether to set a CPU limit at all is genuinely contested. Limits make behaviour predictable and prevent one workload starving others; omitting them lets bursty workloads use idle capacity. For latency-sensitive services on clusters with reasonable headroom, omitting the CPU limit while keeping the request accurate is often the better trade.

The waste that accumulates quietly

Separate from right-sizing, every estate accumulates resources that cost money and do nothing: persistent volumes released when a workload was deleted, load balancers for services that no longer exist, elastic IPs held after a migration, snapshots from a backup process that was replaced, and old container images filling registry storage.

Individually each is small. Collectively, in a cluster that has been running for a few years, it is routinely a meaningful share of the bill, and unlike right-sizing, cleanup carries almost no risk once ownership is confirmed.

Key takeaways

  • Kubernetes allocates on requests, so requests (not usage) determine what you pay.
  • A cluster can be 25% utilised and completely unschedulable.
  • Attribute cost to namespaces and teams before trying to reduce it.
  • Right-size from observed percentiles, and leave real headroom on memory limits.
  • Consider omitting CPU limits on latency-sensitive services while keeping requests accurate.
  • Orphaned volumes, load balancers and snapshots are low-risk, recurring savings.

Frequently asked questions

KubernetesFinOpsCost

See this working on your own infrastructure

A 30-minute walkthrough with a platform engineer. Bring the problem this article describes.