Almost every Kubernetes failure maps to a specific stage of pod creation, and the pod status tells you which one. Pending is scheduling, ImagePullBackOff is the registry, CrashLoopBackOff is the application or its configuration, OOMKilled is the memory limit, and a service with no endpoints is a selector or readiness problem. Read the status first, then the events, then the logs, in that order.
How do you troubleshoot a Kubernetes pod?
Start with the pod status, which identifies the failing stage: Pending means the scheduler could not place it, ImagePullBackOff means the image could not be fetched, CrashLoopBackOff means the container starts and exits, and OOMKilled means it exceeded its memory limit. Then read the pod events for the reason, and only then read the container logs, including the previous container instance if it has already restarted.
Read status, then events, then logs
The most common time-waster is starting with logs. If the container never started, there are no logs, and several minutes go into finding that out. The status tells you which stage failed, the events tell you why, and only then are logs the right place to look.
# 1. Status, which stage failed?
kubectl get pod api-7d9f -o wide
# 2. Events and conditions: why?
kubectl describe pod api-7d9f
# 3. Logs, and the previous instance if it already restarted.
kubectl logs api-7d9f --previous --tail=200
# Cluster-wide events, most recent last.
kubectl get events --sort-by=.lastTimestamp -A | tail -40The five failures you will actually meet
| Status | Means | Most common causes |
|---|---|---|
| Pending | The scheduler could not place the pod | Insufficient allocatable CPU or memory, node taints without tolerations, an unbound PersistentVolumeClaim, affinity rules with no matching node |
| ImagePullBackOff | The image could not be fetched | Wrong tag or digest, missing or wrong imagePullSecret, registry unreachable, rate limiting |
| CrashLoopBackOff | The container starts and exits repeatedly | Application error on startup, missing configuration or secret, failing liveness probe, wrong command or entrypoint |
| OOMKilled | The container exceeded its memory limit | Limit set too low, memory leak, memory scaling with request volume |
| Running but not Ready | Readiness probe failing | Dependency unreachable, probe path wrong, initialDelaySeconds too short for a slow start |
CrashLoopBackOff in detail
This is the most common and the most ambiguous, because it only says the container keeps exiting. The exit code narrows it considerably.
- Exit code 0: the process completed. Usually a command that runs and exits rather than serving, or a misconfigured entrypoint.
- Exit code 1: a generic application error. The logs from the previous instance will normally say what.
- Exit code 137: SIGKILL, almost always OOMKilled. Check the memory limit against observed usage.
- Exit code 143: SIGTERM, terminated deliberately. Often a liveness probe failing and the kubelet restarting it.
- Immediate exit with no logs, usually the entrypoint binary is missing, or the wrong architecture for the node.
Pending pods and why capacity graphs lie
A pending pod with plenty of apparent spare CPU is the classic confusion. The scheduler places pods based on requests, not usage, so a cluster at 30% actual usage can be fully allocated and unable to accept anything.
# The scheduler's reason is in the events, stated plainly.
kubectl describe pod api-7d9f | sed -n '/Events/,$p'
# 0/6 nodes are available: 4 Insufficient memory, 2 node(s) had untolerated taint
# Compare requests against allocatable, not usage against capacity.
kubectl describe node node-3 | sed -n '/Allocated resources/,/^Events/p'Services with no endpoints
When a service returns connection refused, the first check is whether it has any endpoints at all. An empty endpoint list has exactly two causes: the selector matches no pods, or the pods it matches are not Ready.
kubectl get endpoints payments-api
# NAME ENDPOINTS AGE
# payments-api <none> 42d <- selector mismatch or nothing Ready
# Which pods does the selector actually match?
kubectl get pods -l app=payments-api
# And does the service targetPort match the container port?
kubectl get svc payments-api -o yaml | grep -A3 portsA selector typo is the most common cause and the easiest to miss, because both objects look correct in isolation.
When the cluster is the problem
If several unrelated workloads degrade at once, stop investigating the workloads. Check node conditions for MemoryPressure, DiskPressure or NotReady, and check API server latency and error rate. A slow API server makes every controller slow, which surfaces as unrelated-looking failures across the whole cluster.
This is also where a platform view earns its place: correlating a burst of workload failures with a node condition or a control-plane latency change is a glance if the signals share a timeline, and a lengthy cross-referencing exercise if they do not.
Key takeaways
- Status first, events second, logs third: starting with logs wastes time when the container never started.
- Exit code 137 is OOMKilled; check the memory limit against observed usage.
- Always read the previous container instance logs, not the fresh one.
- Pending is a request-versus-allocatable problem, not a usage problem.
- An empty endpoint list means a selector mismatch or nothing Ready, nothing else.
- Several unrelated workloads failing at once points at the node or the control plane.
Frequently asked questions
The container starts, exits, and Kubernetes restarts it with an increasing backoff delay. It indicates the container is failing on or shortly after startup: the exit code and the previous instance logs identify why.
Check that the image reference is correct including the tag or digest, that an imagePullSecret exists and is referenced if the registry is private, that the node can reach the registry, and that you are not being rate limited. The describe output states which of these it is.
The container exceeded its memory limit and was terminated by the kernel with no grace period. Either the limit is too low for real usage, or the application has a leak or a memory profile that scales with load.
The readiness probe is failing. Common causes are a dependency the pod cannot reach, a probe path or port that does not match the application, or an initialDelaySeconds too short for a slow-starting service.
Usually because it has no endpoints: either the selector matches no pods or the matching pods are not Ready. Check the endpoints object first; it distinguishes these immediately.
Use an ephemeral debug container attached to the running pod, which gives you tooling without adding a shell to the production image. This is the reason distroless images remain debuggable.