Cloud Deploy: Promotion, Canary and Automation

Rollout job phases and where they fail, verify jobs that actually gate, canary phases without writing orchestration, automation rules, and a timed rollback.

advanced 24 min lesson hands-on task included

The previous lesson built the objects. This one is what happens between them: the jobs inside a rollout, the gates that can stop it, and the automation that removes the clicks you do not want a human doing.


Topic 1: A Rollout Is a Sequence of Jobs

ONE ROLLOUT IS A SEQUENCE OF JOBS — EACH CAN FAIL AND EACH IS RETRYABLE approve only if requireApproval predeploy a hook: migrate, drain, notify deploy skaffold apply of frozen manifests verify your smoke test, used as a gate postdeploy a hook: warm cache, announce A CANARY REPEATS deploy AND verify PER PHASE canary-25 → verify → canary-50 → verify → stable → verify (a failed verify aborts and rolls back) RETRY, DO NOT RE-RELEASE gcloud deploy rollouts retry-job … A failed verify from a flaky dependency is a retry, not a new release. HOOKS RUN OUTSIDE THE ROLLBACK A predeploy migration has already run when a rollout fails. Same rule as Helm hooks: be backward compatible. WHERE TO LOOK WHEN A ROLLOUT IS STUCK gcloud deploy rollouts describe … → phases[].jobs[].state · then open the Cloud Build log for that job — the real error is there
Five job types per rollout, each independently retryable. The panel on the right is the constraint carried over from Helm hooks — anything a predeploy hook did is still done after a rollback.
gcloud deploy rollouts describe checkout-9f8e7d-to-prod-0001 \
  --delivery-pipeline=checkout --release=checkout-9f8e7d --region=europe-west1 \
  --format='yaml(state, phases)'
phases:
  - id: stable
    state: SUCCEEDED
    deploymentJobs:
      predeployJob:  { state: SUCCEEDED }
      deployJob:     { state: SUCCEEDED }
      verifyJob:     { state: FAILED }        ← this is what stopped it
      postdeployJob: { state: SKIPPED }

Reading a stuck rollout has one reliable path: find the failed job in phases[].jobs[], then open the Cloud Build execution behind it — every job runs as a build, and that log is where the real error is. The rollout resource tells you which job; the build log tells you why.

Hooks are ordinary containers, defined in Skaffold:

# skaffold.yaml
customActions:
  - name: predeploy-migrate
    containers:
      - name: migrate
        image: europe-west1-docker.pkg.dev/acme/apps/migrate:latest
        command: ['/app/migrate']
        args: ['up']
# pipeline stage
    - targetId: prod
      strategy:
        standard:
          predeploy:  { actions: [predeploy-migrate] }
          postdeploy: { actions: [warm-cache] }

The rollback constraint is the same as everywhere else in this course: a predeploy migration has already run when a later job fails. Rolling back returns the manifests; it does not return the schema. Expand-and-contract migrations, or you do not have a rollback — the deployment-strategies lesson states the pattern.


Topic 2: Verify — the Gate That Makes Promotion Mean Something

# skaffold.yaml
verify:
  - name: smoke
    container:
      name: smoke
      image: curlimages/curl:8.8.0
      command: ['sh', '-c']
      args:
        - |
          set -euo pipefail
          BASE="http://checkout.${NAMESPACE}.svc.cluster.local"
          curl -fsS "$BASE/healthz"
          curl -fsS "$BASE/api/v1/status" | grep -q '"database":"ok"'

A failing verify fails the rollout — and inside a canary, aborts the phase and rolls back. That is what turns a smoke test from a report into a gate.

Three levels of verify, in ascending usefulness:

1. it answers            curl /healthz          — catches a broken deploy
2. it works              a real request path    — catches a broken dependency
3. it is not worse       query your metrics     — catches a regression

Level 3 is where a canary earns its keep, and it is the step most teams skip:

verify:
  - name: error-budget
    container:
      name: check
      image: europe-west1-docker.pkg.dev/acme/tools/promquery:1.4
      args:
        - --query=sum(rate(http_requests_total{service="checkout",version="canary",status=~"5.."}[5m]))
                  / sum(rate(http_requests_total{service="checkout",version="canary"}[5m]))
        - --max=0.01
        - --window=5m

A verify job that queries Prometheus, Cloud Monitoring or your own SLO service and exits non-zero on a regression is the difference between a canary and a slow deployment. Verify jobs need enough time to be meaningful — executionTimeout on the target, and a window long enough for the metric to be significant at your traffic volume.


Topic 3: Canary Without Writing Orchestration

      strategy:
        canary:
          runtimeConfig:
            kubernetes:
              gatewayServiceMesh:            # real weighted routing
                httpRoute: checkout
                service: checkout
                deployment: checkout
          canaryDeployment:
            percentages: [5, 25, 50]
            verify: true

Cloud Deploy creates the canary workload, shifts the weight, runs verify between phases, and proceeds or aborts. Two runtime configurations, and the difference matters:

  • serviceNetworking approximates the split with replica counts — no mesh required, and the granularity is limited by how many replicas you run. At 3 replicas, “5%” is not available.
  • gatewayServiceMesh uses Gateway API or Istio for genuine weighted routing, and is what you want for small percentages.

customCanaryDeployment gives per-phase control when the standard phases do not fit — different verify jobs per phase, or a phase that only deploys without shifting traffic.

The preconditions for a canary that is worth the machinery, restated because they decide whether to bother: enough traffic for the metric to be significant, metrics labelled by version, and a written abort condition. Without all three, blue-green with a fast rollback is the better engineering choice, and the strategies lesson makes that case.


Topic 4: Approvals

# On the Target
requireApproval: true
gcloud deploy rollouts approve checkout-9f8e7d-to-prod-0001 \
  --delivery-pipeline=checkout --release=checkout-9f8e7d --region=europe-west1

gcloud deploy rollouts reject  checkout-9f8e7d-to-prod-0001 …

Why this beats an input step in a build:

  • It is an object with a state, not a paused process. Nothing is holding an executor, and a controller restart does not lose it.
  • IAM decides who can approve (roles/clouddeploy.approver), which is a real authorisation boundary rather than a text field.
  • Every approval and rejection is in Cloud Audit Logs, with an identity and a timestamp, retained on your schedule.
  • It outlives the build. Six months later the record still answers “who authorised this”.

Wire notifications to the approval, or it becomes the thing that silently blocks releases:

gcloud pubsub subscriptions create deploy-approvals \
  --topic=clouddeploy-approvals

Cloud Deploy publishes to clouddeploy-operations and clouddeploy-approvals; a small Cloud Run notifier turns those into Slack messages with the approve link.


Topic 5: Automation

AUTOMATION — THE PROMOTIONS YOU DO NOT WANT A HUMAN FOR promoteRelease after a wait, promote to the next target dev → staging, unattended Keeps humans on the gate that matters. advanceRollout advance a canary phase when verify passed no click between 25% and 50% Progressive delivery without a babysitter. repairRollout retry a failed job N times then rollback automatically bounded, not infinite Flaky infrastructure stops paging you. rules: - promoteReleaseRule: name: dev-to-staging wait: 10m selector: [{ target: { id: dev } }] AUTOMATE FORWARD, GATE AT PRODUCTION The shape most teams want: dev and staging promote themselves, production requires an approval, and a repair rule handles transient failures anywhere. Automation needs its own service account and IAM.
Three rule types. The shape most teams converge on is on the right — automate forward through the environments where a human adds nothing, and gate at production.
apiVersion: deploy.cloud.google.com/v1
kind: Automation
metadata:
  name: checkout-flow
serviceAccount: cd-automation@acme.iam.gserviceaccount.com
selector:
  - target: { id: dev }
rules:
  - promoteReleaseRule:
      name: dev-to-staging
      wait: 10m
# Advance a canary phase automatically once verify has passed
  - advanceRolloutRule:
      name: advance-canary
      sourcePhases: [canary-25, canary-50]
      wait: 5m
# Bounded retry, then automatic rollback
  - repairRolloutRule:
      name: repair
      phases: [stable]
      jobs: [deploy]
      repairPhases:
        - retry: { attempts: 2, wait: 60s, backoffMode: BACKOFF_MODE_LINEAR }
        - rollback: { destinationPhase: stable }

repairRolloutRule is the underused one. A deploy job that fails on a transient cluster error currently pages someone; with a bounded retry followed by an automatic rollback, it resolves itself and leaves a record. Bounded is the key word — attempts: 2, not infinite.

Automation needs its own service account with clouddeploy.operator and permission to act on the pipeline. Give it exactly that, since it acts without a human.


Topic 6: Rollback, Retry, and the Drill

Three different recoveries, and choosing correctly matters:

# The job failed for an environmental reason — retry it, do not re-release
gcloud deploy rollouts retry-job checkout-9f8e7d-to-prod-0001 \
  --job-id=deploy --phase-id=stable \
  --delivery-pipeline=checkout --release=checkout-9f8e7d --region=europe-west1

# The release is bad — go back to the previous one
gcloud deploy targets rollback prod \
  --delivery-pipeline=checkout --region=europe-west1

# Roll back to a specific release rather than the immediately previous one
gcloud deploy targets rollback prod --release=checkout-3c4d5e \
  --delivery-pipeline=checkout --region=europe-west1

Rollback creates a new rollout of an existing release. The manifests were rendered and frozen when that release was created, so there is no rebuild, no re-render and no dependency on the source repository still being in that state. That is what makes it fast, and it is worth stating because it is the property most home-grown deploy scripts lack.

What a rollback does not undo, and this belongs in your runbook:

· a predeploy or postdeploy hook that already ran
· a schema migration
· data written by the new version
· messages consumed, emails sent, webhooks delivered
· anything outside the target — a DNS change, a feature flag

The drill, quarterly, in staging first and then production:

1. Note the current release. Start the clock at the DECISION, not the command.
2. gcloud deploy targets rollback
3. Watch the rollout reach SUCCEEDED
4. Verify traffic is on the old version — check the workload, not the console
5. Record the total. That number is your real recovery time.

Try it yourself: make the verify job fail on a canary at 25% and watch Cloud Deploy abort and roll back without anyone intervening. That single observation is the whole argument for putting the smoke test in verify rather than in a build step after the deploy.

Common mistake: setting verify: true on a canary with a verify job that only checks /healthz. The new version answers, the canary advances through every phase, and a regression in error rate — the exact thing a canary exists to catch — passes every gate. The check has to compare the canary against the stable version on a metric that would move; anything less is ceremony with extra waiting.