Kubernetes

Multi-cluster Kubernetes management: patterns that hold up

Why organisations end up with many Kubernetes clusters, what actually gets hard at that point, and the operational patterns that survive contact with a growing fleet.

Harshit SengarFounder and platform lead, DevOpsArkPublished 4 March 2026 · Updated 2 August 202610 min read
TL;DR

Multi-cluster is usually accidental rather than designed, and the problems it creates are not technical so much as informational: nobody knows how many clusters exist, what versions they run, who owns them or what they cost. The patterns that hold up are a single inventory derived from live state, one delivery path with per-cluster configuration, policy defined once and evaluated everywhere, and upgrades planned from a deprecated-API impact list rather than attempted and observed.

Short answer

What is multi-cluster Kubernetes management?

Multi-cluster Kubernetes management is the practice of operating several Kubernetes clusters as one estate rather than individually: maintaining a single inventory, a shared deployment path, one policy baseline and consistent observability across clusters that may span providers, regions and environments.

Nobody plans to have twelve clusters

Multi-cluster estates are rarely designed. They accumulate. One cluster per environment becomes three. A regulated workload needs isolation, so that is four. A team in another region needs lower latency, five. An acquisition arrives with its own, six. Someone spins up a cluster for a proof of concept and it becomes load-bearing, seven. By the time anyone counts, the number is larger than the last person to ask believed.

There are good reasons for many of these: blast radius isolation, data residency, regional latency, hard multi-tenancy boundaries. The problem is not the number of clusters. It is that the operational model was designed for one.

What actually gets hard

Running twelve clusters is not twelve times harder than running one. Some things scale linearly, some barely change, and a few become qualitatively different problems.

ConcernAt one clusterAt twelve
InventoryYou remember itNobody knows the full list without querying
Version skewOne versionA range spanning several minor versions, with the oldest unsupported
PolicyApplied onceApplied inconsistently, drifting apart
UpgradesA planned afternoonA programme nobody wants to own
CostOne bill lineDistributed across accounts with no per-team view
AccessOne RBAC setTwelve RBAC sets, reviewed as zero
Incident responseYou know where things runThe first question is which cluster

The pattern is consistent: the technical operations stay manageable, and the informational ones collapse. The binding constraint on a large fleet is almost always knowing rather than doing.

Pattern one: one inventory, derived from live state

The foundational pattern is a single inventory of clusters, nodes, namespaces and workloads that is derived from the clusters themselves rather than maintained by hand. Any inventory that requires a human to update it is wrong within a month, and a wrong inventory is worse than none because people trust it.

Derived inventory also changes what questions are answerable. "Which clusters are more than two minor versions behind", "which namespaces run privileged containers", "what does the payments team cost across every environment": these are one query against a fleet index and an afternoon of work without one.

Pattern two: one delivery path, per-cluster configuration

The mistake here is copying deployment configuration per cluster. It works for three and fails for twelve, because a change to the shared parts has to be made twelve times and will be made eleven times.

The pattern that holds is a single delivery definition with cluster-specific configuration held separately: the same image digest promoted everywhere, differing only in the values each environment supplies. Rollout strategy and approval requirements are properties of the target rather than of the artifact, so production can require a canary with a metric gate while a development cluster takes the change immediately.

# One definition, many targets. Configuration differs; the artifact does not.
artifact:
  image: ghcr.io/ark/api
  digest: sha256:8f2c1ad...        # promoted unchanged

targets:
  - cluster: dev-eu-west
    strategy: rolling
    approval: none
  - cluster: staging-eu-west
    strategy: canary
    steps: [25, 100]
    approval: none
  - cluster: prod-eu-west
    strategy: canary
    steps: [5, 25, 50, 100]
    gate:
      successRate: ">= 99.5%"
      p99Latency: "<= 400ms"
    approval: required

Pattern three: policy defined once, evaluated everywhere

Cluster baselines drift apart the moment they are maintained per cluster. Pod security standards, network policy defaults, image provenance requirements and RBAC boundaries should be defined once and evaluated against every cluster continuously, with the differences reported rather than assumed absent.

The useful output is not a pass or fail per cluster but a list of specific deviations: this namespace in this cluster runs privileged containers, these three clusters have no default deny network policy. That is a work list. A compliance percentage is not.

Pattern four: upgrades planned from impact, not attempted and observed

Version skew is the most reliable symptom of a fleet that has outgrown its operating model. Clusters fall behind because nobody can predict what an upgrade will break, so the upgrade is deferred, and the longer it is deferred the less predictable it becomes.

Breaking the cycle requires the impact list before the upgrade: which workloads in which clusters use APIs the target version removes, and where in the manifests to change them. With that list, an upgrade is a scoped piece of work. Without it, it is an experiment performed on production.

Upgrade in waves. Development clusters first, then a low-traffic production cluster, then the rest. Verify workload health between waves and halt on failure. The point of the fleet is that you have somewhere safe to go first.

How DevOpsArk handles multi-cluster

DevOpsArk connects clusters agentlessly through the Kubernetes API, builds one continuously derived inventory across providers, drives delivery through a single path with per-cluster strategy and approvals, evaluates one policy set against every cluster, and produces deprecated-API impact lists from the workloads actually running before an upgrade.

Because cost attribution and security posture read the same inventory, a question like "what does this team spend across all clusters and what is its worst exposure" is one query rather than a project.

Key takeaways

  • Multi-cluster estates accumulate rather than being designed, and the operating model usually lags behind.
  • What breaks at scale is informational (inventory, versions, ownership, cost), not the mechanics of running clusters.
  • Derive the inventory from live state; a hand-maintained one is wrong within a month.
  • Use one delivery definition with per-cluster configuration rather than copied manifests.
  • Define policy once and report specific deviations, not a compliance percentage.
  • Plan upgrades from a deprecated-API impact list, and roll them out in waves.

Frequently asked questions

KubernetesMulti-clusterPlatform engineering

See this working on your own infrastructure

A 30-minute walkthrough with a platform engineer. Bring the problem this article describes.