Kubernetes is a set of controllers that continuously compare desired state against actual state and act on the difference. The API server is the only thing that talks to etcd; everything else talks to the API server. Understanding that one loop explains most cluster behaviour, including why changes are eventually consistent and why a slow API server makes everything look broken at once.
What is Kubernetes architecture?
Kubernetes architecture consists of a control plane (the API server, etcd, scheduler and controller manager), and worker nodes running a kubelet and container runtime. The control plane stores desired state and runs controllers that continuously reconcile actual state towards it, while the kubelet on each node makes the local reality match what was assigned to it.
One idea explains most of it
Kubernetes is a collection of control loops. Each loop watches a kind of object, compares the state that object declares against the state that exists, and takes an action to reduce the difference. That is the whole model. Deployments, ReplicaSets, services, ingresses, persistent volumes and autoscalers are all the same pattern applied to different nouns.
Two consequences follow immediately, and they explain a great deal of day-to-day behaviour. Changes are eventually consistent, so an apply that returns successfully means the desired state was recorded, not that it has happened. And there is no single component whose failure stops everything. Instead, a degraded control plane causes reconciliation to slow down, which surfaces as unrelated-looking symptoms all over the cluster.
The control plane
| Component | Responsibility | What its failure looks like |
|---|---|---|
| kube-apiserver | The only front door. Validates, authorises and persists every change | Everything appears broken at once; kubectl hangs; controllers stop reconciling |
| etcd | The datastore holding all cluster state | API server errors or high latency; writes rejected when the database is out of space |
| kube-scheduler | Assigns pods to nodes based on requests, affinity, taints and topology | New pods sit in Pending; existing pods are unaffected |
| kube-controller-manager | Runs the built-in reconciliation loops | Objects stop converging: ReplicaSets do not scale, endpoints go stale |
| cloud-controller-manager | Provisions provider resources such as load balancers and volumes | Services of type LoadBalancer stay pending; volumes fail to attach |
The node
A node runs three things that matter operationally. The kubelet takes the pods assigned to that node and makes them exist, reporting status back. The container runtime (containerd in most clusters now) actually starts and stops containers. kube-proxy programmes the network rules that make service addresses route somewhere.
The kubelet is also the component that enforces resource limits and runs probes, which makes it the source of the two most common workload failures: a container OOMKilled for exceeding its memory limit, and a container restarted for failing a liveness probe. Neither is a scheduling problem, and looking at the scheduler for them wastes time.
Where a pod actually comes from
Tracing one deployment through the system is the fastest way to make the architecture concrete.
- You apply a Deployment. The API server validates it, authorises you, and writes it to etcd.
- The Deployment controller notices a Deployment with no matching ReplicaSet and creates one.
- The ReplicaSet controller notices it has zero pods and needs three, and creates three Pod objects with no node assigned.
- The scheduler notices three unassigned pods, evaluates requests, affinity, taints and topology, and writes a node name onto each.
- The kubelet on each node notices a pod assigned to it, pulls the image, and starts the container.
- The endpoints controller notices new ready pods matching a service selector and adds them to its endpoints.
- kube-proxy on every node updates its rules so traffic to the service address reaches the new pods.
Every failure mode maps to a step. Pending means step four did not complete. ImagePullBackOff means step five failed. A service with no endpoints means step six has nothing ready to add. Knowing which step stalled is most of the diagnosis.
Key takeaways
- Kubernetes is control loops reconciling desired state against actual state. That one idea explains most behaviour.
- Only the API server talks to etcd, which makes API server latency the highest-value control-plane metric.
- A successful apply means the desired state was recorded, not that it has happened.
- The kubelet enforces limits and runs probes, so OOMKills and probe restarts are node-level, not scheduling, problems.
- Each pod-creation step maps to a specific failure mode, which makes diagnosis a matter of finding the stalled step.
Frequently asked questions
The control plane is the set of components that store cluster state and drive reconciliation: the API server, etcd, the scheduler, the controller manager and, in cloud environments, the cloud controller manager. It decides what should exist; nodes make it exist.
The kubelet runs on every node, takes the pods assigned to that node, and makes them real: pulling images, starting containers, running liveness and readiness probes, enforcing resource limits and reporting status back to the API server.
etcd is the key-value store holding all cluster state: every object you have created and its current status. Only the API server talks to it directly, and backing it up is what makes cluster state recoverable.
Running workloads keep running, because the kubelet on each node continues managing the pods it already has. What stops is change: no new deployments, no rescheduling of failed pods, no scaling and no service endpoint updates.
A Pod is one or more containers scheduled together. A ReplicaSet keeps a specified number of identical pods running. A Deployment manages ReplicaSets to provide rolling updates and rollback. You almost always create a Deployment and let it create the rest.