Kubernetes overspend is almost always the gap between what workloads request and what they use, plus idle node capacity nobody owns. Fixing it needs attribution first, because you cannot optimise what you cannot assign, then concrete right-sizing values rather than percentage advice, then cleanup of the orphaned resources that accumulate quietly.
Why is Kubernetes so expensive?
Kubernetes costs are usually driven by the gap between resource requests and actual usage. Because the scheduler reserves capacity based on requests, over-provisioned workloads consume cluster capacity they never use, forcing more nodes than the real workload requires. Idle node capacity, orphaned volumes and oversized node pools account for most of the rest.
Requests are the bill, not usage
The central fact of Kubernetes cost is that the scheduler allocates on requests, not on usage. A pod requesting 2 CPU and using 0.1 occupies 2 CPU of schedulable capacity. Ten of those and you are paying for a node to do almost nothing.
This is why cluster utilisation graphs mislead so consistently. A cluster showing 25% CPU usage and refusing to schedule new pods is not a paradox: it is fully allocated and barely used.
Attribution comes before optimisation
A cluster has one invoice line and many tenants. Until node cost is distributed across namespaces and workloads, nobody owns any part of it, and shared costs that nobody owns do not get reduced.
Allocation works by distributing each node cost across the pods scheduled onto it, weighted by requests and usage. The important detail is what to do with the unallocated remainder, the capacity that is paid for and not requested by anything. Hiding it in a shared bucket removes the incentive to reduce it; distributing it proportionally makes idle capacity visible to the teams whose over-requesting caused the node to exist.
The changes that produce real savings
| Change | Typical impact | Effort |
|---|---|---|
| Right-size requests against observed usage | Large, usually the biggest single item | Low per workload, needs data |
| Remove orphaned volumes, load balancers and IPs | Moderate and recurring | Low |
| Consolidate under-utilised node pools | Moderate | Medium, needs a scheduling review |
| Use spot or preemptible capacity for tolerant workloads | Large where applicable | Medium, needs disruption handling |
| Scale non-production to zero outside working hours | Moderate | Low |
| Commitment coverage for steady-state usage | Moderate | Low, but needs confidence in the baseline |
| Reduce over-replication of low-traffic services | Small individually, adds up | Low |
Right-sizing without causing an incident
The failure mode of enthusiastic right-sizing is setting limits from average usage and then having pods OOMKilled during a peak. Two rules avoid it.
- Set requests near the sustained usage percentile you actually observe, commonly p90 over several weeks, including the busiest day of the week.
- Set memory limits with genuine headroom above peak, because exceeding a memory limit is a hard kill with no grace period. CPU limits are more forgiving, since throttling degrades rather than terminates.
# Observed over 30 days: memory p50 340Mi, p99 620Mi, peak 710Mi.
# CPU p50 0.08, p99 0.31.
resources:
requests:
memory: "700Mi" # near sustained peak, this is what gets scheduled
cpu: "100m" # near p50; bursting is fine for CPU
limits:
memory: "1Gi" # headroom above observed peak; exceeding this kills
# No CPU limit: throttling a latency-sensitive service to save nothing
# is a common and avoidable self-inflicted problem.Whether to set a CPU limit at all is genuinely contested. Limits make behaviour predictable and prevent one workload starving others; omitting them lets bursty workloads use idle capacity. For latency-sensitive services on clusters with reasonable headroom, omitting the CPU limit while keeping the request accurate is often the better trade.
The waste that accumulates quietly
Separate from right-sizing, every estate accumulates resources that cost money and do nothing: persistent volumes released when a workload was deleted, load balancers for services that no longer exist, elastic IPs held after a migration, snapshots from a backup process that was replaced, and old container images filling registry storage.
Individually each is small. Collectively, in a cluster that has been running for a few years, it is routinely a meaningful share of the bill, and unlike right-sizing, cleanup carries almost no risk once ownership is confirmed.
Key takeaways
- Kubernetes allocates on requests, so requests (not usage) determine what you pay.
- A cluster can be 25% utilised and completely unschedulable.
- Attribute cost to namespaces and teams before trying to reduce it.
- Right-size from observed percentiles, and leave real headroom on memory limits.
- Consider omitting CPU limits on latency-sensitive services while keeping requests accurate.
- Orphaned volumes, load balancers and snapshots are low-risk, recurring savings.
Frequently asked questions
Distribute each node cost across the pods on it, weighted by requests and usage, then map workloads to owning teams from a service catalogue. Distribute unallocated node capacity proportionally rather than hiding it, so idle spend has an owner.
A request is what the scheduler reserves and what determines placement and cost. A limit is the ceiling the kernel enforces at run time. Exceeding a memory limit kills the container immediately; exceeding a CPU limit throttles it.
It is genuinely debated. Limits make behaviour predictable and protect neighbours; omitting them lets bursty workloads use spare capacity and avoids throttling latency-sensitive services for no benefit. Accurate requests matter more than either choice.
It helps, but horizontal autoscaling with over-provisioned requests scales the over-provisioning too. Fix the requests first; autoscaling then operates on accurate numbers.
It depends entirely on the starting point. Estates where requests were never revisited after initial deployment tend to have the largest gap, and right-sizing plus orphan cleanup usually accounts for most of the achievable saving.