A rolling update is the default and is fine for changes you are confident about. Blue/green gives you an instant switch and an instant revert at the cost of running two full environments. Canary limits exposure by sending a small share of real traffic to the new version first, and it is the only strategy that catches problems your tests did not. None of them help unless something is measuring release health and can revert without a human.
What is a canary deployment?
A canary deployment releases a new version to a small share of production traffic while the rest continues to the current version. Its health is measured against defined metrics such as success rate and latency, and the share is increased in steps only if those metrics stay within threshold. If they do not, the canary is withdrawn and only the small share was ever exposed.
The three strategies, honestly compared
| Rolling | Blue/green | Canary | |
|---|---|---|---|
| Extra capacity needed | A little (surge) | Double, briefly | A little |
| Exposure during rollout | Increasing share | None, then everyone | Small, controlled share |
| Revert speed | Another rolling update | Instant traffic switch | Instant, withdraw the canary |
| Catches problems tests missed | Partially | No, all-or-nothing switch | Yes, this is the point |
| Complexity | Built in | Moderate: two environments, one switch | Highest: needs traffic splitting and metrics |
| Database migration friendly | Requires compatible schema | Hard, two versions on one database | Requires compatible schema |
The honest summary: rolling is the default because it is free, blue/green buys revert speed with money, and canary buys real-world validation with complexity. Most organisations should use different strategies in different environments rather than picking one.
Rolling updates and the mistake everyone makes
Kubernetes rolling updates replace pods gradually, controlled by maxUnavailable and maxSurge. The common mistake is treating the rollout as complete when the pods are Ready, because readiness usually means "the process started and answered a health check", not "this version works".
A readiness probe that returns 200 from a handler which does nothing tells you the HTTP server is up. It does not tell you the new version can reach its database, that its configuration parsed, or that its p99 latency is acceptable. Rollouts that succeed while the service breaks are almost always this.
strategy:
type: RollingUpdate
rollingUpdate:
maxUnavailable: 0 # never drop below the desired count
maxSurge: 1 # add one at a time
# A readiness probe worth having checks dependencies, not just liveness.
readinessProbe:
httpGet:
path: /readyz # verifies database and cache connectivity
port: 8080
initialDelaySeconds: 5
periodSeconds: 5
failureThreshold: 3Blue/green: paying for revert speed
Blue/green runs the new version as a complete parallel environment, verifies it, then switches traffic in one step. Revert is the same switch in reverse, which is as fast as a deployment gets. The costs are the doubled capacity during the overlap and the fact that any shared stateful dependency (most obviously the database) is not duplicated, so schema compatibility is still your problem.
It suits releases where the switch itself must be atomic: a coordinated frontend and backend change, or a release where a partially updated fleet would be incoherent.
Canary: the only one that finds unknown problems
A canary sends a small share of real traffic to the new version and judges it on real behaviour. This is the only strategy that catches problems your test suite did not model, which is the category most production incidents come from: an unusual request shape, a slow path only real data triggers, a memory profile that only appears under sustained load.
The essential part is the gate. A canary without automated evaluation is just a slow rolling update with extra steps, because the failure mode it is supposed to prevent (nobody notices for twenty minutes) is exactly what it does not fix.
What to gate on
- Success rate of the canary compared with the stable version over the same window, not against an absolute number.
- Latency at a high percentile (p95 or p99), because a mean hides the regression you care about.
- Error log rate for patterns that are new since the canary started.
- Saturation signals such as memory approaching the limit, which catch a leak before it becomes an OOMKill.
Rollback is the feature, not the strategy
The strategy determines how exposure grows. What determines whether a bad release becomes an incident is whether something reverts it automatically. A team with rolling updates and automatic rollback recovers faster than a team with canaries that a human has to evaluate.
ArkCD treats this as the primary property: you nominate the signals that define a healthy release, the rollout evaluates them at each step, and a breached gate reverts to the last known-good revision on its own and records the gate result that caused it. The strategy is then a per-environment choice rather than a safety mechanism.
Key takeaways
- Choose a strategy per environment; there is no single correct answer for the whole estate.
- Readiness means the process started, not that the version works. Probe dependencies, not just the port.
- Blue/green buys instant revert with doubled capacity, and does nothing about your database schema.
- Canary is the only strategy that finds problems your tests did not model.
- Gate a canary against the stable version rather than an absolute threshold.
- Automatic rollback matters more than the choice of strategy.
Frequently asked questions
Blue/green runs two complete environments and switches all traffic at once, so exposure goes from zero to everyone. Canary shifts a small share of traffic to the new version first and increases it in steps based on measured health, so exposure is always bounded.
Progressive delivery is the practice of releasing a change to an increasing share of users while automatically evaluating its health, and halting or reverting if it degrades. Canary deployment is its most common form.
Long enough to see representative traffic through the affected code path. For a high-volume service that can be minutes; for an endpoint used a few times an hour it may be a day. Time is a poor proxy: the useful measure is whether enough requests have exercised the change.
Yes. Replica-count weighting gives approximate traffic splitting with plain Kubernetes services, and most ingress controllers support weighted routing. A mesh gives finer control and better per-version telemetry, but it is not a prerequisite.
They constrain all of them, because two application versions share one database. The standard approach is expand and contract: deploy a schema change that both versions tolerate, roll out the application, then remove the old columns in a later release once nothing reads them.