The previous lesson built the objects. This one is what happens between them: the jobs inside a rollout, the gates that can stop it, and the automation that removes the clicks you do not want a human doing.
Topic 1: A Rollout Is a Sequence of Jobs
gcloud deploy rollouts describe checkout-9f8e7d-to-prod-0001 \
--delivery-pipeline=checkout --release=checkout-9f8e7d --region=europe-west1 \
--format='yaml(state, phases)'
phases:
- id: stable
state: SUCCEEDED
deploymentJobs:
predeployJob: { state: SUCCEEDED }
deployJob: { state: SUCCEEDED }
verifyJob: { state: FAILED } ← this is what stopped it
postdeployJob: { state: SKIPPED }
Reading a stuck rollout has one reliable path: find the failed job in phases[].jobs[], then open the Cloud Build execution behind it — every job runs as a build, and that log is where the real error is. The rollout resource tells you which job; the build log tells you why.
Hooks are ordinary containers, defined in Skaffold:
# skaffold.yaml
customActions:
- name: predeploy-migrate
containers:
- name: migrate
image: europe-west1-docker.pkg.dev/acme/apps/migrate:latest
command: ['/app/migrate']
args: ['up']
# pipeline stage
- targetId: prod
strategy:
standard:
predeploy: { actions: [predeploy-migrate] }
postdeploy: { actions: [warm-cache] }
The rollback constraint is the same as everywhere else in this course: a predeploy migration has already run when a later job fails. Rolling back returns the manifests; it does not return the schema. Expand-and-contract migrations, or you do not have a rollback — the deployment-strategies lesson states the pattern.
Topic 2: Verify — the Gate That Makes Promotion Mean Something
# skaffold.yaml
verify:
- name: smoke
container:
name: smoke
image: curlimages/curl:8.8.0
command: ['sh', '-c']
args:
- |
set -euo pipefail
BASE="http://checkout.${NAMESPACE}.svc.cluster.local"
curl -fsS "$BASE/healthz"
curl -fsS "$BASE/api/v1/status" | grep -q '"database":"ok"'
A failing verify fails the rollout — and inside a canary, aborts the phase and rolls back. That is what turns a smoke test from a report into a gate.
Three levels of verify, in ascending usefulness:
1. it answers curl /healthz — catches a broken deploy
2. it works a real request path — catches a broken dependency
3. it is not worse query your metrics — catches a regression
Level 3 is where a canary earns its keep, and it is the step most teams skip:
verify:
- name: error-budget
container:
name: check
image: europe-west1-docker.pkg.dev/acme/tools/promquery:1.4
args:
- --query=sum(rate(http_requests_total{service="checkout",version="canary",status=~"5.."}[5m]))
/ sum(rate(http_requests_total{service="checkout",version="canary"}[5m]))
- --max=0.01
- --window=5m
A verify job that queries Prometheus, Cloud Monitoring or your own SLO service and exits non-zero on a regression is the difference between a canary and a slow deployment. Verify jobs need enough time to be meaningful — executionTimeout on the target, and a window long enough for the metric to be significant at your traffic volume.
Topic 3: Canary Without Writing Orchestration
strategy:
canary:
runtimeConfig:
kubernetes:
gatewayServiceMesh: # real weighted routing
httpRoute: checkout
service: checkout
deployment: checkout
canaryDeployment:
percentages: [5, 25, 50]
verify: true
Cloud Deploy creates the canary workload, shifts the weight, runs verify between phases, and proceeds or aborts. Two runtime configurations, and the difference matters:
serviceNetworkingapproximates the split with replica counts — no mesh required, and the granularity is limited by how many replicas you run. At 3 replicas, “5%” is not available.gatewayServiceMeshuses Gateway API or Istio for genuine weighted routing, and is what you want for small percentages.
customCanaryDeployment gives per-phase control when the standard phases do not fit — different verify jobs per phase, or a phase that only deploys without shifting traffic.
The preconditions for a canary that is worth the machinery, restated because they decide whether to bother: enough traffic for the metric to be significant, metrics labelled by version, and a written abort condition. Without all three, blue-green with a fast rollback is the better engineering choice, and the strategies lesson makes that case.
Topic 4: Approvals
# On the Target
requireApproval: true
gcloud deploy rollouts approve checkout-9f8e7d-to-prod-0001 \
--delivery-pipeline=checkout --release=checkout-9f8e7d --region=europe-west1
gcloud deploy rollouts reject checkout-9f8e7d-to-prod-0001 …
Why this beats an input step in a build:
- It is an object with a state, not a paused process. Nothing is holding an executor, and a controller restart does not lose it.
- IAM decides who can approve (
roles/clouddeploy.approver), which is a real authorisation boundary rather than a text field. - Every approval and rejection is in Cloud Audit Logs, with an identity and a timestamp, retained on your schedule.
- It outlives the build. Six months later the record still answers “who authorised this”.
Wire notifications to the approval, or it becomes the thing that silently blocks releases:
gcloud pubsub subscriptions create deploy-approvals \
--topic=clouddeploy-approvals
Cloud Deploy publishes to clouddeploy-operations and clouddeploy-approvals; a small Cloud Run notifier turns those into Slack messages with the approve link.
Topic 5: Automation
apiVersion: deploy.cloud.google.com/v1
kind: Automation
metadata:
name: checkout-flow
serviceAccount: cd-automation@acme.iam.gserviceaccount.com
selector:
- target: { id: dev }
rules:
- promoteReleaseRule:
name: dev-to-staging
wait: 10m
# Advance a canary phase automatically once verify has passed
- advanceRolloutRule:
name: advance-canary
sourcePhases: [canary-25, canary-50]
wait: 5m
# Bounded retry, then automatic rollback
- repairRolloutRule:
name: repair
phases: [stable]
jobs: [deploy]
repairPhases:
- retry: { attempts: 2, wait: 60s, backoffMode: BACKOFF_MODE_LINEAR }
- rollback: { destinationPhase: stable }
repairRolloutRule is the underused one. A deploy job that fails on a transient cluster error currently pages someone; with a bounded retry followed by an automatic rollback, it resolves itself and leaves a record. Bounded is the key word — attempts: 2, not infinite.
Automation needs its own service account with clouddeploy.operator and permission to act on the pipeline. Give it exactly that, since it acts without a human.
Topic 6: Rollback, Retry, and the Drill
Three different recoveries, and choosing correctly matters:
# The job failed for an environmental reason — retry it, do not re-release
gcloud deploy rollouts retry-job checkout-9f8e7d-to-prod-0001 \
--job-id=deploy --phase-id=stable \
--delivery-pipeline=checkout --release=checkout-9f8e7d --region=europe-west1
# The release is bad — go back to the previous one
gcloud deploy targets rollback prod \
--delivery-pipeline=checkout --region=europe-west1
# Roll back to a specific release rather than the immediately previous one
gcloud deploy targets rollback prod --release=checkout-3c4d5e \
--delivery-pipeline=checkout --region=europe-west1
Rollback creates a new rollout of an existing release. The manifests were rendered and frozen when that release was created, so there is no rebuild, no re-render and no dependency on the source repository still being in that state. That is what makes it fast, and it is worth stating because it is the property most home-grown deploy scripts lack.
What a rollback does not undo, and this belongs in your runbook:
· a predeploy or postdeploy hook that already ran
· a schema migration
· data written by the new version
· messages consumed, emails sent, webhooks delivered
· anything outside the target — a DNS change, a feature flag
The drill, quarterly, in staging first and then production:
1. Note the current release. Start the clock at the DECISION, not the command.
2. gcloud deploy targets rollback
3. Watch the rollout reach SUCCEEDED
4. Verify traffic is on the old version — check the workload, not the console
5. Record the total. That number is your real recovery time.
Try it yourself: make the verify job fail on a canary at 25% and watch Cloud Deploy abort and roll back without anyone intervening. That single observation is the whole argument for putting the smoke test in verify rather than in a build step after the deploy.
Common mistake: setting verify: true on a canary with a verify job that only checks /healthz. The new version answers, the canary advances through every phase, and a regression in error rate — the exact thing a canary exists to catch — passes every gate. The check has to compare the canary against the stable version on a metric that would move; anything less is ceremony with extra waiting.