Every GCP resource is zonal, regional or multi-regional, and that classification decides what survives what. Getting it written down is most of a disaster recovery plan; the rest is two numbers.
Topic 1: Scope Decides What Dies Together
Zonal — dies with its zone. A GCE VM, a zonal persistent disk, a zonal GKE cluster’s control plane, Filestore Basic.
Regional — survives a zone. A regional MIG, a regional persistent disk (synchronously replicated across two zones), Cloud SQL with REGIONAL availability, a regional GKE cluster, regional GCS.
Multi-regional — survives a region. Multi-region GCS, Spanner, the global load balancer, Cloud DNS, Artifact Registry multi-region.
# Verify rather than assume — these are the checks that find surprises
gcloud compute instances list --format='table(name, zone, status)'
gcloud compute disks list --format='table(name, zone, region, type)'
gcloud sql instances list --format='table(name, region, gceZone, availabilityType)'
gcloud container clusters list --format='table(name, location, locationType)'
The misfiled cases, which are the ones that cause incidents:
- A zonal persistent disk attached to a VM in a regional MIG. The MIG is regional; the disk is not, and it cannot be attached in another zone.
- A zonal GKE cluster. The nodes may span zones while the control plane does not — a zone failure takes the API server with it.
- Cloud SQL without
REGIONAL. A single-zone database under a multi-zone application tier. - Regional GCS presented as durable — it is, against disk failure, not against losing the region.
Topic 2: RTO and RPO Before Anything Else
RTO — how long until service is restored. RPO — how much data you may lose, measured in time.
They are set by different mechanisms, and confusing them produces designs that are expensive and still wrong:
RPO is a function of replication and backup frequency.
RTO is a function of how fast capacity comes up and traffic moves.
Continuous replication buys a near-zero RPO with a slow RTO. A warm standby buys a fast RTO even with hourly backups. You choose them separately, and you choose them per workload — a single organisation-wide number is always wrong somewhere.
Get the numbers in writing from whoever owns the revenue, because the cost curve is steep and engineers asked to “make it resilient” reliably build too much or too little.
Topic 3: The Four Strategies, on GCP
| Strategy | RTO | What it is on GCP |
|---|---|---|
| Backup & restore | Hours–days | Backups in multi-region GCS, Terraform to rebuild, restore on demand |
| Pilot light | Tens of minutes | Cross-region Cloud SQL replica, images in Artifact Registry, MIG at size 0 |
| Warm standby | Minutes | A scaled-down live stack in region B, already behind the global LB |
| Active/active | Near zero | Both regions serving; Spanner, or data partitioned by region |
The global load balancer changes the shape of this, and it is worth being explicit about why. With one anycast IP and backends in two regions, regional failover is a health-check outcome — no DNS change, no TTL to wait out, no client-side caching to fight. That single property removes the most fragile part of most DR plans.
gcloud compute backend-services add-backend checkout-be --global \
--instance-group=checkout-mig-euw1 --instance-group-region=europe-west1 \
--balancing-mode=RATE --max-rate-per-instance=100
gcloud compute backend-services add-backend checkout-be --global \
--instance-group=checkout-mig-euw4 --instance-group-region=europe-west4 \
--balancing-mode=RATE --max-rate-per-instance=100
With that in place, losing europe-west1 means requests land in europe-west4 within a health-check interval. The remaining problem is entirely the data layer — which is where the real work is.
Topic 4: The Data Layer Is the Hard Part
| Data | Cross-region option | RPO |
|---|---|---|
| Cloud SQL | Cross-region read replica, manual promotion | Seconds–minutes |
| AlloyDB | Cross-region replica | Seconds |
| Spanner | Multi-region configuration, synchronous | Zero |
| Firestore | Multi-region location | Near zero |
| GCS | Multi-region bucket, or turbo replication | Seconds–minutes |
| Persistent Disk | Scheduled snapshots, or Asynchronous Replication | Minutes |
| Artifact Registry | Multi-region, or explicit copies | n/a |
| Secret Manager | Multi-region replication policy | Seconds |
The rows people forget are the last three. A DR region with a database and no container images, no secrets and no TLS certificate is not a DR region — it is a subnet. Every drill finds at least one of them.
Spanner is the only zero-RPO multi-region write option, and that is what you are paying for. Everything else means choosing between asynchronous replication (some data loss on failover) and a single-region write path (no loss, no cross-region writes).
Topic 5: Static Stability
The principle that separates designs that survive a large event from ones that fail during recovery: a recovery that depends on creating new resources depends on the control plane, and the control plane is what degrades in a large event.
Concretely, on GCP:
- Capacity already running beats capacity you intend to create. Warm standby over pilot light, if the RTO is tight.
- Quotas are per project, per region. The DR region’s quota is usually the default, because nothing has ever run there. Raise it in advance — a quota increase during an incident is a support case.
- Reservations hold capacity in a zone for you, and are the only way to guarantee that the instances you plan to launch can be launched.
- Avoid recovery steps that need a human with unique access. A runbook naming a person is a single point of failure with a holiday schedule.
Topic 6: Drills, and the Runbook They Produce
An untested plan does not have a slow RTO — it has an unknown one.
Monthly Restore one backup into a scratch project. Verify the data,
not just that the command exited zero.
Quarterly Zone failure. Delete the instances in one zone of a regional
MIG, or use a firewall rule to isolate it, and watch the service.
Annually Full regional failover, announced, with the runbook open and
somebody recording timings for every step.
Every drill produces two artifacts: an updated runbook and a list of what surprised you. Real findings from real drills, all of them GCP-specific:
- The DR region had no container images in Artifact Registry.
- The CMEK key was regional and did not exist in the second region, so the restored disks could not be decrypted.
- The service account existed but had no IAM bindings in the DR project.
- The managed certificate had never been provisioned for the failover hostname.
- The quota for the machine type was the default 8 vCPUs.
Write the runbook for the tired version of yourself: exact commands, exact resource names, explicit decision points, and who tells customers what.
Try it yourself: list every disk in your project with gcloud compute disks list and check the zone column against where the workload claims to be resilient. A zonal disk under a regional workload is extremely common and completely invisible until the zone goes.
Common mistake: measuring RTO from “the restore finished” rather than from “we decided to fail over”. The restore is often the fastest part — the time goes into noticing, deciding, finding the right backup, discovering the application needs a new connection name, and getting agreement to switch. Time the whole thing, decision included, or the number in your plan is fiction.