Deployment Strategies

Recreate, rolling, blue-green and canary compared on blast radius and cost — plus the database constraint that caps how progressive your delivery can be.

advanced 20 min lesson hands-on task included

Every deployment strategy is a way of buying a smaller blast radius. The question is what you are willing to pay in capacity, complexity and time — and whether your data layer permits two versions to run at once.


Topic 1: The Four Strategies

FOUR STRATEGIES — YOU ARE BUYING A SMALLER BLAST RADIUS RECREATE cheapest stop all, start all downtime by design no version overlap Batch jobs, or anything that canno t run two versions. ROLLING free replace pods gradually both versions live the Kubernetes default The sane default when versions are compatible. BLUE / GREEN 2× capacity full second environment switch traffic at once instant rollback When rollback speed matters more t han cost. CANARY small overhead 1% → 25% → 100% metrics decide each step auto rollback on error The strongest option, and it needs real SLOs. A CANARY WITHOUT METRICS IS A SLOW OUTAGE If nothing automatically compares error rate and latency between versions, you have added steps and no safety. THE DATABASE DECIDES YOUR CEILING Two versions live at once means the schema must serve both. Expand-and-contract, or no progressive delivery. MEASURE ROLLBACK, NOT DEPLOYMENT The number that matters is how long from "this is wrong" to "traffic is on the old version" — and the only way to know it is to do it.
Read left to right as increasing safety and increasing cost. The two panels at the bottom are the constraints that decide whether the right-hand options are available to you at all.

Recreate — stop everything, start the new version. Downtime by design, no version overlap. Correct for a batch job, a singleton that cannot run twice, or a schema change that genuinely requires exclusive access.

Rolling — replace instances gradually. Both versions serve simultaneously. Free, built into Kubernetes Deployments, and the sane default when versions are compatible.

strategy:
  type: RollingUpdate
  rollingUpdate:
    maxSurge: 25%
    maxUnavailable: 0     # never go below the current capacity

maxUnavailable: 0 is the setting worth defaulting to: it costs a little extra capacity during the rollout and guarantees you never dip below your current replica count.

Blue-green — a complete second environment; switch traffic at once. Instant rollback because the old version is still running. Costs double capacity for the duration, and doubles nothing else — which is why it is the simplest way to get a fast rollback.

Canary — send a small fraction of traffic to the new version, compare, then widen. The strongest option, and the one that requires real signals to be worth anything.

DowntimeExtra capacityRollback speedNeeds
RecreateYesNoneRedeploy oldNothing
RollingNo~25%Rolling backCompatible versions
Blue-greenNo100%SecondsTraffic switch
CanaryNoSmallSecondsTraffic split and metrics

Topic 2: Rolling Updates, and the Probe That Makes Them Safe

A rolling update is only safe if Kubernetes can tell a healthy pod from a started one. That is the readiness probe’s job, and a wrong probe turns a rolling update into a rolling outage.

readinessProbe:
  httpGet: { path: /healthz, port: 8080 }
  initialDelaySeconds: 5
  periodSeconds: 5
  failureThreshold: 3

Two rules from the Kubernetes module that matter most here:

  • The readiness probe must not check dependencies. A probe that fails when the database is briefly slow marks every pod unready at once, and the rollout stalls with no capacity.
  • maxUnavailable: 0 plus a working readiness probe means the new pod serves traffic only when it is genuinely ready, and the old one is removed only then.

minReadySeconds is underused: it makes a pod wait before counting as available, which catches the process that starts, passes one probe and then crashes.


Topic 3: Blue-Green

       ┌─────────┐
       │ Service │  selector: version=blue
       └────┬────┘
    ┌───────┴───────┐
 [blue: v1.5]   [green: v1.6]   ← both running, only blue receives traffic

The switch is one label change:

kubectl patch service checkout -p '{"spec":{"selector":{"version":"green"}}}'

What it buys: rollback is the same command in reverse, and it takes seconds because the old version never stopped running.

What to get right:

  • Keep blue running for long enough to be confident — an hour, a day. That is the window in which rollback is instant.
  • Sessions and connections. Existing long-lived connections stay on blue; be explicit about whether that is acceptable.
  • Anything stateful. Two versions writing to one database is the constraint in Topic 5, and blue-green does not avoid it — it just makes the switch instant.
  • Cost. Double capacity for the overlap. On a large service that is a real number, and it is usually still cheaper than an outage.

Topic 4: Canary, and Why Most Canaries Are Decorative

# A weighted split with the Gateway API
apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
spec:
  rules:
    - backendRefs:
        - { name: checkout-stable, port: 80, weight: 90 }
        - { name: checkout-canary, port: 80, weight: 10 }

Progressive delivery controllers — Argo Rollouts, Flagger — automate the widening and the abort:

# Argo Rollouts, with an analysis step that can actually fail
strategy:
  canary:
    steps:
      - setWeight: 10
      - pause: { duration: 5m }
      - analysis:
          templates: [{ templateName: error-rate }]
      - setWeight: 50
      - pause: { duration: 5m }
      - analysis:
          templates: [{ templateName: error-rate }]
      - setWeight: 100
# The analysis is the part that matters
apiVersion: argoproj.io/v1alpha1
kind: AnalysisTemplate
metadata: { name: error-rate }
spec:
  metrics:
    - name: error-rate
      interval: 1m
      failureLimit: 2
      provider:
        prometheus:
          address: http://prometheus.monitoring:9090
          query: |
            sum(rate(http_requests_total{service="checkout",version="canary",status=~"5.."}[2m]))
            /
            sum(rate(http_requests_total{service="checkout",version="canary"}[2m]))

A canary without an automated comparison is a slow deployment with extra steps. If nothing queries error rate, latency or a business metric and fails the rollout, then a human is watching a dashboard — which works until the deploy happens at 4pm on a Friday.

Three practical requirements for a canary that earns its keep:

  • Enough traffic. At 10% of 20 requests per minute, statistical significance takes hours. Low-traffic services should use blue-green.
  • Metrics labelled by version. If your dashboards cannot separate canary from stable, the comparison is impossible.
  • A defined abort condition, written down: which metric, what threshold, over what window.

Topic 5: The Database Decides Your Ceiling

Every strategy except recreate runs two versions of your code simultaneously. Both versions talk to the same database, which means the schema must serve both — and that constraint, not the deployment tooling, is what caps how progressive your delivery can be.

Expand and contract, over several releases:

Release N     add the new column, nullable. Old code ignores it.
Release N+1   write to both old and new columns. Backfill existing rows.
Release N+2   read from the new column only.
Release N+3   drop the old column.

At every step, the previous release still works — which is what makes rollback real. A migration that renames or drops a column in one release means the previous version is broken the moment it runs, and you have given up rollback for that release whether or not anyone noticed.

The same applies to message formats, cache keys and API contracts. Two versions in flight means every interface between them must be tolerant in both directions. The Helm module covers the hook-and-rollback side of this; the principle is identical.


Topic 6: Measure Rollback, Not Deployment

The number that matters is time from “this is wrong” to “traffic is on the old version”, and the only way to know it is to do it.

□ How long does a rollback take, measured, this quarter?
□ Who can perform it — one person, or anyone on call?
□ Does it require a rebuild? (It should not.)
□ Does it require a human to notice, or does a check trigger it?
□ What does it NOT undo — migrations, consumed messages, sent emails?

That last question is the one teams skip. A rollback returns the code; it does not un-send an email, un-charge a card or un-consume a queue message. Knowing which of those apply is what separates a rollback plan from a rollback command.

Practise it deliberately. A monthly rollback drill in staging — including the decision, not just the command — is the cheapest reliability work available, and it is what turns the number in the table above from an estimate into a fact.

Try it yourself: deploy a version that fails 10% of requests behind a canary at 10% weight. Whether your setup catches it tells you if your canary is real. If it does not, the missing piece is a query, and writing that query is the highest-value hour in this lesson.

Common mistake: adopting canary deployments because they are the sophisticated option, on a service with 30 requests per minute and no per-version metrics. The rollout takes 40 minutes, nothing can be concluded from the data, and the extra machinery makes deploys rarer — which makes each one bigger and riskier. Rolling with maxUnavailable: 0 and a fast, practised rollback beats a decorative canary every time.