Multi-cluster is usually accidental rather than designed, and the problems it creates are not technical so much as informational: nobody knows how many clusters exist, what versions they run, who owns them or what they cost. The patterns that hold up are a single inventory derived from live state, one delivery path with per-cluster configuration, policy defined once and evaluated everywhere, and upgrades planned from a deprecated-API impact list rather than attempted and observed.
What is multi-cluster Kubernetes management?
Multi-cluster Kubernetes management is the practice of operating several Kubernetes clusters as one estate rather than individually: maintaining a single inventory, a shared deployment path, one policy baseline and consistent observability across clusters that may span providers, regions and environments.
Nobody plans to have twelve clusters
Multi-cluster estates are rarely designed. They accumulate. One cluster per environment becomes three. A regulated workload needs isolation, so that is four. A team in another region needs lower latency, five. An acquisition arrives with its own, six. Someone spins up a cluster for a proof of concept and it becomes load-bearing, seven. By the time anyone counts, the number is larger than the last person to ask believed.
There are good reasons for many of these: blast radius isolation, data residency, regional latency, hard multi-tenancy boundaries. The problem is not the number of clusters. It is that the operational model was designed for one.
What actually gets hard
Running twelve clusters is not twelve times harder than running one. Some things scale linearly, some barely change, and a few become qualitatively different problems.
| Concern | At one cluster | At twelve |
|---|---|---|
| Inventory | You remember it | Nobody knows the full list without querying |
| Version skew | One version | A range spanning several minor versions, with the oldest unsupported |
| Policy | Applied once | Applied inconsistently, drifting apart |
| Upgrades | A planned afternoon | A programme nobody wants to own |
| Cost | One bill line | Distributed across accounts with no per-team view |
| Access | One RBAC set | Twelve RBAC sets, reviewed as zero |
| Incident response | You know where things run | The first question is which cluster |
The pattern is consistent: the technical operations stay manageable, and the informational ones collapse. The binding constraint on a large fleet is almost always knowing rather than doing.
Pattern one: one inventory, derived from live state
The foundational pattern is a single inventory of clusters, nodes, namespaces and workloads that is derived from the clusters themselves rather than maintained by hand. Any inventory that requires a human to update it is wrong within a month, and a wrong inventory is worse than none because people trust it.
Derived inventory also changes what questions are answerable. "Which clusters are more than two minor versions behind", "which namespaces run privileged containers", "what does the payments team cost across every environment": these are one query against a fleet index and an afternoon of work without one.
Pattern two: one delivery path, per-cluster configuration
The mistake here is copying deployment configuration per cluster. It works for three and fails for twelve, because a change to the shared parts has to be made twelve times and will be made eleven times.
The pattern that holds is a single delivery definition with cluster-specific configuration held separately: the same image digest promoted everywhere, differing only in the values each environment supplies. Rollout strategy and approval requirements are properties of the target rather than of the artifact, so production can require a canary with a metric gate while a development cluster takes the change immediately.
# One definition, many targets. Configuration differs; the artifact does not.
artifact:
image: ghcr.io/ark/api
digest: sha256:8f2c1ad... # promoted unchanged
targets:
- cluster: dev-eu-west
strategy: rolling
approval: none
- cluster: staging-eu-west
strategy: canary
steps: [25, 100]
approval: none
- cluster: prod-eu-west
strategy: canary
steps: [5, 25, 50, 100]
gate:
successRate: ">= 99.5%"
p99Latency: "<= 400ms"
approval: requiredPattern three: policy defined once, evaluated everywhere
Cluster baselines drift apart the moment they are maintained per cluster. Pod security standards, network policy defaults, image provenance requirements and RBAC boundaries should be defined once and evaluated against every cluster continuously, with the differences reported rather than assumed absent.
The useful output is not a pass or fail per cluster but a list of specific deviations: this namespace in this cluster runs privileged containers, these three clusters have no default deny network policy. That is a work list. A compliance percentage is not.
Pattern four: upgrades planned from impact, not attempted and observed
Version skew is the most reliable symptom of a fleet that has outgrown its operating model. Clusters fall behind because nobody can predict what an upgrade will break, so the upgrade is deferred, and the longer it is deferred the less predictable it becomes.
Breaking the cycle requires the impact list before the upgrade: which workloads in which clusters use APIs the target version removes, and where in the manifests to change them. With that list, an upgrade is a scoped piece of work. Without it, it is an experiment performed on production.
How DevOpsArk handles multi-cluster
DevOpsArk connects clusters agentlessly through the Kubernetes API, builds one continuously derived inventory across providers, drives delivery through a single path with per-cluster strategy and approvals, evaluates one policy set against every cluster, and produces deprecated-API impact lists from the workloads actually running before an upgrade.
Because cost attribution and security posture read the same inventory, a question like "what does this team spend across all clusters and what is its worst exposure" is one query rather than a project.
Key takeaways
- Multi-cluster estates accumulate rather than being designed, and the operating model usually lags behind.
- What breaks at scale is informational (inventory, versions, ownership, cost), not the mechanics of running clusters.
- Derive the inventory from live state; a hand-maintained one is wrong within a month.
- Use one delivery definition with per-cluster configuration rather than copied manifests.
- Define policy once and report specific deviations, not a compliance percentage.
- Plan upgrades from a deprecated-API impact list, and roll them out in waves.
Frequently asked questions
As few as your isolation requirements allow. Legitimate reasons to add a cluster are blast radius separation, data residency, regional latency and hard tenancy boundaries. Team preference and environment separation can often be handled with namespaces and policy instead.
Fewer, larger clusters give better bin packing and less operational overhead per workload. More, smaller clusters give better blast radius isolation and simpler tenancy. Most organisations end up in the middle, and the deciding factor is usually regulatory or regional rather than technical.
Cluster sprawl is the accumulation of clusters beyond what anyone is actively managing: clusters created for a project, a demo or a team that outlive their purpose. The symptom is being unable to state the full list from memory, and the cost is unpatched, unmonitored, unowned infrastructure.
Track version and support status per cluster in one place, generate deprecated-API impact reports before upgrading, and upgrade in waves starting with the lowest-risk cluster. Consistency comes from making upgrades routine, not from mandating a version.
Yes, provided it works through the Kubernetes API rather than provider-specific interfaces. DevOpsArk connects EKS, AKS, GKE, OpenShift and self-managed clusters through the same path and presents them in one inventory.