Resilience and Disaster Recovery on GCP

Zonal, regional and multi-regional resource scope, turning availability into an RTO and an RPO, and why a global load balancer makes regional failover a health-check outcome.

advanced 24 min lesson hands-on task included

Every GCP resource is zonal, regional or multi-regional, and that classification decides what survives what. Getting it written down is most of a disaster recovery plan; the rest is two numbers.


Topic 1: Scope Decides What Dies Together

EVERY GCP RESOURCE IS ZONAL, REGIONAL, OR MULTI-REGIONAL ZONAL dies with the zone GCE VM · zonal PD zonal GKE · Filestore basic REGIONAL survives a zone regional MIG · regional PD Cloud SQL HA · regional GKE MULTI-REGIONAL survives a region GCS multi-region · Spanner global LB · Cloud DNS AND THE FOUR STRATEGIES BUILT ON THEM Backup & restore RTO hours–days GCS backups + Terraform to rebuild Pilot light RTO tens of minutes data replicating, MIG at size 0 Warm standby RTO minutes a small live copy behind the global LB Active/active RTO near zero both regions serving; Spanner or partitioned data A global load balancer makes regional failover a health-check outcome rather than a DNS change — which is the largest single DR advantage GCP offers.
The top row is the classification; the bottom is what you can build on it. The banner at the foot is GCP's genuine structural advantage over the alternatives.

Zonal — dies with its zone. A GCE VM, a zonal persistent disk, a zonal GKE cluster’s control plane, Filestore Basic.

Regional — survives a zone. A regional MIG, a regional persistent disk (synchronously replicated across two zones), Cloud SQL with REGIONAL availability, a regional GKE cluster, regional GCS.

Multi-regional — survives a region. Multi-region GCS, Spanner, the global load balancer, Cloud DNS, Artifact Registry multi-region.

# Verify rather than assume — these are the checks that find surprises
gcloud compute instances list --format='table(name, zone, status)'
gcloud compute disks list --format='table(name, zone, region, type)'
gcloud sql instances list --format='table(name, region, gceZone, availabilityType)'
gcloud container clusters list --format='table(name, location, locationType)'

The misfiled cases, which are the ones that cause incidents:

  • A zonal persistent disk attached to a VM in a regional MIG. The MIG is regional; the disk is not, and it cannot be attached in another zone.
  • A zonal GKE cluster. The nodes may span zones while the control plane does not — a zone failure takes the API server with it.
  • Cloud SQL without REGIONAL. A single-zone database under a multi-zone application tier.
  • Regional GCS presented as durable — it is, against disk failure, not against losing the region.

Topic 2: RTO and RPO Before Anything Else

RTO — how long until service is restored. RPO — how much data you may lose, measured in time.

They are set by different mechanisms, and confusing them produces designs that are expensive and still wrong:

RPO is a function of replication and backup frequency.
RTO is a function of how fast capacity comes up and traffic moves.

Continuous replication buys a near-zero RPO with a slow RTO. A warm standby buys a fast RTO even with hourly backups. You choose them separately, and you choose them per workload — a single organisation-wide number is always wrong somewhere.

Get the numbers in writing from whoever owns the revenue, because the cost curve is steep and engineers asked to “make it resilient” reliably build too much or too little.


Topic 3: The Four Strategies, on GCP

StrategyRTOWhat it is on GCP
Backup & restoreHours–daysBackups in multi-region GCS, Terraform to rebuild, restore on demand
Pilot lightTens of minutesCross-region Cloud SQL replica, images in Artifact Registry, MIG at size 0
Warm standbyMinutesA scaled-down live stack in region B, already behind the global LB
Active/activeNear zeroBoth regions serving; Spanner, or data partitioned by region

The global load balancer changes the shape of this, and it is worth being explicit about why. With one anycast IP and backends in two regions, regional failover is a health-check outcome — no DNS change, no TTL to wait out, no client-side caching to fight. That single property removes the most fragile part of most DR plans.

gcloud compute backend-services add-backend checkout-be --global \
  --instance-group=checkout-mig-euw1 --instance-group-region=europe-west1 \
  --balancing-mode=RATE --max-rate-per-instance=100

gcloud compute backend-services add-backend checkout-be --global \
  --instance-group=checkout-mig-euw4 --instance-group-region=europe-west4 \
  --balancing-mode=RATE --max-rate-per-instance=100

With that in place, losing europe-west1 means requests land in europe-west4 within a health-check interval. The remaining problem is entirely the data layer — which is where the real work is.


Topic 4: The Data Layer Is the Hard Part

DataCross-region optionRPO
Cloud SQLCross-region read replica, manual promotionSeconds–minutes
AlloyDBCross-region replicaSeconds
SpannerMulti-region configuration, synchronousZero
FirestoreMulti-region locationNear zero
GCSMulti-region bucket, or turbo replicationSeconds–minutes
Persistent DiskScheduled snapshots, or Asynchronous ReplicationMinutes
Artifact RegistryMulti-region, or explicit copiesn/a
Secret ManagerMulti-region replication policySeconds

The rows people forget are the last three. A DR region with a database and no container images, no secrets and no TLS certificate is not a DR region — it is a subnet. Every drill finds at least one of them.

Spanner is the only zero-RPO multi-region write option, and that is what you are paying for. Everything else means choosing between asynchronous replication (some data loss on failover) and a single-region write path (no loss, no cross-region writes).


Topic 5: Static Stability

The principle that separates designs that survive a large event from ones that fail during recovery: a recovery that depends on creating new resources depends on the control plane, and the control plane is what degrades in a large event.

Concretely, on GCP:

  • Capacity already running beats capacity you intend to create. Warm standby over pilot light, if the RTO is tight.
  • Quotas are per project, per region. The DR region’s quota is usually the default, because nothing has ever run there. Raise it in advance — a quota increase during an incident is a support case.
  • Reservations hold capacity in a zone for you, and are the only way to guarantee that the instances you plan to launch can be launched.
  • Avoid recovery steps that need a human with unique access. A runbook naming a person is a single point of failure with a holiday schedule.

Topic 6: Drills, and the Runbook They Produce

An untested plan does not have a slow RTO — it has an unknown one.

Monthly    Restore one backup into a scratch project. Verify the data,
           not just that the command exited zero.

Quarterly  Zone failure. Delete the instances in one zone of a regional
           MIG, or use a firewall rule to isolate it, and watch the service.

Annually   Full regional failover, announced, with the runbook open and
           somebody recording timings for every step.

Every drill produces two artifacts: an updated runbook and a list of what surprised you. Real findings from real drills, all of them GCP-specific:

  • The DR region had no container images in Artifact Registry.
  • The CMEK key was regional and did not exist in the second region, so the restored disks could not be decrypted.
  • The service account existed but had no IAM bindings in the DR project.
  • The managed certificate had never been provisioned for the failover hostname.
  • The quota for the machine type was the default 8 vCPUs.

Write the runbook for the tired version of yourself: exact commands, exact resource names, explicit decision points, and who tells customers what.

Try it yourself: list every disk in your project with gcloud compute disks list and check the zone column against where the workload claims to be resilient. A zonal disk under a regional workload is extremely common and completely invisible until the zone goes.

Common mistake: measuring RTO from “the restore finished” rather than from “we decided to fail over”. The restore is often the fastest part — the time goes into noticing, deciding, finding the right backup, discovering the application needs a new connection name, and getting agreement to switch. Time the whole thing, decision included, or the number in your plan is fiction.