Most Kubernetes monitoring collects far more than it uses. The signals that predict real failure are a short list: pod restart patterns, CPU throttling against limits, memory usage against limits, pending pods, node conditions and control-plane latency. Alert on those against behavioural baselines rather than fixed thresholds, and treat everything else as diagnostic context you look at after an alert rather than something you page on.
What is Kubernetes monitoring?
Kubernetes monitoring is the collection and analysis of cluster, node, workload and pod signals to determine whether a cluster and the applications on it are healthy. It covers control-plane health, node conditions and capacity, pod scheduling and restarts, and resource usage measured against the requests and limits each workload declares.
Why Kubernetes monitoring is different
Monitoring a fleet of virtual machines is mostly a question of whether each machine is healthy. Kubernetes breaks that assumption in two ways. First, the unit you care about (the workload) is not the unit that runs (the pod), and pods are deliberately disposable. A pod dying is not an incident; a pod dying repeatedly is. Second, the scheduler sits between the workload and the hardware, so a workload can be perfectly healthy and still fail to run because the cluster has nowhere to put it.
That means Kubernetes monitoring has to answer three separate questions rather than one: is the control plane working, does the cluster have somewhere to run things, and are the workloads themselves healthy. Teams that only instrument the third are surprised by the first two.
The signals that actually predict failure
A short list does most of the work. Everything else is diagnostic context, worth collecting, not worth paging on.
| Signal | What it tells you | Why it matters |
|---|---|---|
| Pod restart rate | A container is exiting and being restarted | A restart loop is the single most common visible failure, and the pattern distinguishes a crash from an OOMKill from a failing probe |
| Memory usage against limit | How close a container is to being killed | Exceeding a memory limit is a hard kill with no grace period, and it is the most common cause of mystery restarts |
| CPU throttling | The container is being slowed by its CPU limit | Throttling produces latency without any error, so it is invisible unless you measure it directly |
| Pending pods | The scheduler cannot place a pod | Almost always capacity, taints or a missing volume, and it means a deployment is not actually complete |
| Node conditions | MemoryPressure, DiskPressure, PIDPressure, NotReady | Node-level pressure evicts pods, so it becomes many workload incidents at once |
| API server latency and error rate | Control-plane health | A slow API server makes every controller slow, which looks like unrelated failures everywhere |
Requests, limits and why they distort your metrics
Kubernetes resource metrics are meaningless without the requests and limits they are measured against. A container at 90% CPU is fine if its limit is generous and terrible if it is being throttled. A node at 40% actual CPU usage can be completely unschedulable because the pods on it have reserved everything through their requests.
This produces the most common piece of Kubernetes monitoring confusion: a cluster that looks half idle on a usage graph while nothing new can be scheduled. Requests are a reservation; usage is consumption. You need both, and you need to know which one you are looking at.
- Track usage as a ratio of the limit, because that ratio predicts throttling and OOMKills.
- Track requests as a ratio of node allocatable capacity, because that ratio predicts scheduling failures.
- Track the gap between request and usage, because that gap is what you are paying for and not using.
Alerting without drowning
The failure mode of Kubernetes alerting is volume. One node going unhealthy can produce an alert per pod on that node, per service those pods back, and per synthetic check that touches those services. Sixty notifications describing one event is worse than one, because the responder now has to do the correlation the system should have done.
Three rules that reduce volume without reducing coverage
- Correlate before notifying. Signals sharing a root (the same node, the same deployment, the same dependency) should collapse into one incident with contributing signals listed inside it.
- Baseline instead of threshold. A fixed number cannot be right for both a Monday peak and a Sunday trough. Alert on deviation from the workload own history for that hour and day.
- Suppress the expected. A rolling update causes restarts by design. If the platform knows a rollout is in flight, the restarts it causes should not page anyone.
The measurable outcome to aim for is not "fewer alerts" but a higher proportion of alerts that lead to an action. Track how often each rule fires and how often anyone does anything about it, and delete the rules that produce work without producing decisions.
What to collect but not alert on
Plenty of data is valuable during an investigation and useless as a page. Kubernetes events, container filesystem usage, network throughput per pod, image pull durations, scheduler latency and kubelet queue depth all fall into this category. Collect them, retain them long enough to look backwards, and surface them next to the alert rather than as alerts of their own.
Retention deserves a deliberate decision. A live-only view answers "is it broken now". Capacity planning, seasonality and post-incident review all need history, and the question "was it like this last Tuesday" is one of the most useful an on-call engineer can ask.
How DevOpsArk approaches this
DevOpsArk collects cluster and workload metrics through the Kubernetes API without an in-cluster agent, baselines each series against its own history including weekly seasonality, and draws deployments on the same timeline as the metrics so the most common investigation (did this start when we shipped something) is a glance rather than a cross-system comparison.
Alerts are correlated into incidents before they are sent, routed by the ownership recorded against each service, and enriched with the recent deployments, matching log patterns and prior occurrences before the notification arrives.
Key takeaways
- Monitor three layers separately: control plane, cluster capacity, and workload health.
- Resource metrics are meaningless without the requests and limits they are measured against.
- A cluster can look idle on usage graphs and be completely unschedulable on requests.
- Alert on deviation from a workload own baseline, not on a fixed threshold guessed at launch.
- Correlate related signals into one incident before notifying anyone.
- Measure the proportion of alerts that lead to action, and delete the rules that never do.
Frequently asked questions
Pod restart rate, memory usage against limit, CPU throttling, pending pod count, node conditions such as MemoryPressure and NotReady, and API server latency and error rate. These six predict most real failures; the rest is diagnostic context.
A container was killed because it exceeded its memory limit. There is no grace period and no warning. The kernel terminates the process. The usual causes are a limit set too low for real traffic, a memory leak, or a workload whose memory scales with request volume in a way nobody modelled.
When a container reaches its CPU limit, the kernel restricts it rather than killing it. The application still works but becomes slower, which shows up as latency with no errors. It is invisible unless you specifically measure throttled time.
The scheduler cannot place them. The common causes are insufficient allocatable CPU or memory across nodes, node taints without matching tolerations, a persistent volume claim that cannot be bound, or node affinity rules with no matching node.
No. Prometheus is the most common collector and works well, but the Kubernetes metrics API provides resource metrics directly, and platforms such as DevOpsArk can read them through the cluster API without any in-cluster component.
Long enough for the questions you actually ask. Live-only retention cannot support capacity planning, seasonality analysis or post-incident review. A common arrangement is full resolution for a few weeks and downsampled data for a year or more.