Everything in this module, assembled once, in code, and then deliberately broken. The build is the easy half; the drills are what turn it into something you can defend in a review.
Topic 1: The Target
The requirements, stated as constraints rather than as a resource list:
1. Survives the loss of any one zone, with evidence.
2. No VM, database or container has a public IP — enforced by org policy.
3. No downloadable service-account key exists anywhere in the org.
4. The data tier cannot reach the internet at all, and you can prove it.
5. Data at rest uses a customer-managed key you can disable.
6. Every resource is labelled well enough to attribute its cost.
7. `terraform destroy` leaves nothing behind — no orphaned disks,
snapshots, buckets, addresses or keys.
Constraints 2 and 3 are org policies, not conventions. Constraint 7 is not housekeeping — an environment that cannot be destroyed cleanly cannot be rebuilt reliably, and rebuild-ability is what constraint 1 depends on.
Topic 2: Phase 1 — Organisation and Projects
organization/
├── folders/prod org policies: no external IP, eu-only, no SA keys
│ ├── host-project Shared VPC owner — the network lives here
│ ├── svc-app service project: GKE or Cloud Run
│ └── svc-data service project: Cloud SQL, GCS
├── folders/nonprod the same shape, weaker quotas
└── folders/shared Artifact Registry, Cloud Build, logging sink target
Decisions to write down, because these are what a reviewer asks about:
- Which constraints you set at the organisation, which at the folder, and why nothing is at the project.
- Why the network lives in a host project and the workloads in service projects — and what a service project team can and cannot change.
- Which projects have a budget, and what happens when it is exceeded.
- Where audit logs are routed, and who can delete them (nobody, ideally).
# Verify the guardrails are actually in effect on a child project
gcloud resource-manager org-policies describe compute.vmExternalIpAccess \
--project=svc-app --effective
gcloud resource-manager org-policies describe iam.disableServiceAccountKeyCreation \
--project=svc-app --effective
Topic 3: Phase 2 — Network and Front Door
Shared VPC (custom mode), europe-west1, three zones
subnet: app 10.20.0.0/20 PGA on · secondary ranges for GKE pods/services
subnet: data 10.20.32.0/22 PGA on · NO default route to the internet
Cloud NAT reserved static IPs, logging on, app subnet only
PSC endpoint → Cloud SQL, in the data subnet
firewall default deny egress; allow by network tag, one reason per rule
Cloud DNS private zone for internal names
flow logs on, sampled, exported to BigQuery
Global external ALB
Google-managed certificate
Cloud Armor: preconfigured WAF rules in preview, plus a rate-based rule
health check + the 35.191.0.0/16 and 130.211.0.0/22 firewall rule
backends in two regions, RATE balancing with a measured max
The three things most likely to be wrong, and a reviewer should check them first: the health-check firewall rule (its absence makes everything 502), Private Google Access on both subnets (its absence makes private VMs hang), and whether the data subnet genuinely has no egress route rather than merely no NAT.
Topic 4: Phase 3 — Workload, Data and Keys
Workload (choose one and justify it)
Cloud Run ingress=internal-and-cloud-load-balancing, min-instances≥1,
deployed by digest, secrets from Secret Manager
or GKE regional cluster, Workload Identity on, private nodes,
release channel + maintenance window set
Data
Cloud SQL REGIONAL availability, private IP only, PITR on,
CMEK, deletion protection, IAM database auth
GCS uniform bucket-level access, public access prevention,
CMEK, lifecycle rule with a written retention reason
Secret Manager per-secret IAM, user-managed replication in-region
Keys
key ring in the same region as the data
90-day rotation, destruction delay left long
service agents granted encrypterDecrypter — and nobody else
Prove the identity story rather than assuming it: there should be no service-account key file anywhere, the workload should reach the database through IAM authentication, and every secret access should appear in the audit log with the workload’s own identity.
Topic 5: Phase 4 — The Seven Drills
Each drill: establish the state, run it, verify with a command, record the time.
1. Kill a zone. Delete the instances in one zone of the regional MIG, or isolate it with a firewall rule. Expect: the service stays up, the load balancer stops routing there, capacity is replaced. Watch for: a zonal disk, a zonal cluster, or a Cloud SQL instance that turned out to be ZONAL.
2. Violate an org policy. Try to create a VM with an external IP. Expect: a constraint violation naming the policy. Verify: the message says constraints/…, not PERMISSION_DENIED — and you can explain the difference.
3. Restore the database with PITR. To five minutes before a deliberate bad UPDATE. Expect: a new instance with a new connection name. Record: the time from decision to a served query.
4. Rotate a CMEK key version. Create a new version, then disable the old one and observe what breaks and what does not. Expect: existing data still readable (it uses the old version) until you disable it — at which point it is not. Record: what recovered when you re-enabled.
5. Prove the data tier cannot egress. From a data-subnet host via IAP, attempt an outbound connection. Expect: a timeout. Verify: the flow log entry showing no path. This is the artifact that answers an auditor.
6. Break a health check. Make the backend return 500 on /healthz. Expect: backends go UNHEALTHY and the LB returns 502 with failed_to_pick_backend. Record: how long from first failure to first 502, and how long to recover.
7. Promote a cross-region replica. Expect: a manual promotion that permanently breaks replication, plus an application that needs repointing. Record: the full elapsed time, and write down what you would have to change to halve it.
Topic 6: What to Produce, and How It Is Judged
Four artifacts:
- The Terraform repository — modules with real interfaces, a GCS backend with state locking, no hardcoded project IDs, and a clean
planagainst the deployed state. - An architecture note — one page, a decision and a rejected alternative for each significant choice. “Shared VPC rather than peering, because peering is not transitive and we will have five service projects within a year.”
- A runbook — the seven drills with commands, verification and measured timings, written for someone who has not seen the system.
- A cost breakdown — monthly cost by line item with the three largest identified. If you cannot explain the load balancer forwarding-rule charge or the inter-zone egress, that is the gap to close.
The questions you should be able to answer without notes:
- Which single resource, if deleted, causes the largest outage?
- What happens if the CMEK key is disabled? If it is destroyed?
- Which of your resources are zonal, and what fails with each zone?
- Can a service project team change the network? Should they be able to?
- Where would a
PERMISSION_DENIEDcome from in this design — IAM, org policy, or VPC-SC? - What in this build breaks first if traffic increases tenfold?
Common mistake: building the environment, screenshotting the architecture, and never breaking it. Drills 1, 4, 5 and 6 all fail on landing zones that look perfect in Terraform — the zonal disk under a regional workload, the CMEK key in the wrong region, the data subnet that turned out to have a NAT route, and the missing health-check firewall rule. The build proves it exists; the drills prove somebody can operate it.