Project: A Production Landing Zone on GCP

Build the whole module as one environment — org policies, Shared VPC, private workloads, CMEK, a global load balancer — then break it seven ways and prove it recovers.

advanced 45 min lesson hands-on task included

Everything in this module, assembled once, in code, and then deliberately broken. The build is the easy half; the drills are what turn it into something you can defend in a review.


Topic 1: The Target

PROJECT TARGET — EVERY BOX IS SOMETHING YOU MUST BE ABLE TO DEFEND ORGANIZATION — org policies, essential contacts, billing export folder: prod folder: nonprod folder: shared folder: sandbox HOST PROJECT — Shared VPC, custom mode, 3 zones subnet: app (private) PGA on · Cloud NAT egress secondary ranges for GKE subnet: data (private) PSC endpoint → Cloud SQL no route to the internet global external ALB + Cloud Armor + managed certificate firewall: deny-all egress by default, allow by tag PROVE IT □ kill a zone, stay up□ deny an org policy violation□ restore the DB with PITR□ rotate a CMEK key version□ block egress, read the log□ fail a health check → 502□ promote a cross-region replica SERVICE PROJECTS ATTACH TO THE HOST — THEY NEVER OWN A NETWORK One team, one service project, one budget, and a Shared VPC subnet they can use but not change. That separation is the whole point of the layout. Build it with Terraform, not the console — the console version cannot be reviewed, cannot be diffed, and cannot be destroyed cleanly at the end of the exercise.
Every box maps to a lesson in this module. If you cannot say which lesson explains a box and why the alternative was rejected, that is the lesson to re-read before building it.

The requirements, stated as constraints rather than as a resource list:

1. Survives the loss of any one zone, with evidence.
2. No VM, database or container has a public IP — enforced by org policy.
3. No downloadable service-account key exists anywhere in the org.
4. The data tier cannot reach the internet at all, and you can prove it.
5. Data at rest uses a customer-managed key you can disable.
6. Every resource is labelled well enough to attribute its cost.
7. `terraform destroy` leaves nothing behind — no orphaned disks,
   snapshots, buckets, addresses or keys.

Constraints 2 and 3 are org policies, not conventions. Constraint 7 is not housekeeping — an environment that cannot be destroyed cleanly cannot be rebuilt reliably, and rebuild-ability is what constraint 1 depends on.


Topic 2: Phase 1 — Organisation and Projects

organization/
├── folders/prod         org policies: no external IP, eu-only, no SA keys
│   ├── host-project     Shared VPC owner — the network lives here
│   ├── svc-app          service project: GKE or Cloud Run
│   └── svc-data         service project: Cloud SQL, GCS
├── folders/nonprod      the same shape, weaker quotas
└── folders/shared       Artifact Registry, Cloud Build, logging sink target

Decisions to write down, because these are what a reviewer asks about:

  • Which constraints you set at the organisation, which at the folder, and why nothing is at the project.
  • Why the network lives in a host project and the workloads in service projects — and what a service project team can and cannot change.
  • Which projects have a budget, and what happens when it is exceeded.
  • Where audit logs are routed, and who can delete them (nobody, ideally).
# Verify the guardrails are actually in effect on a child project
gcloud resource-manager org-policies describe compute.vmExternalIpAccess \
  --project=svc-app --effective
gcloud resource-manager org-policies describe iam.disableServiceAccountKeyCreation \
  --project=svc-app --effective

Topic 3: Phase 2 — Network and Front Door

Shared VPC (custom mode), europe-west1, three zones
  subnet: app   10.20.0.0/20   PGA on · secondary ranges for GKE pods/services
  subnet: data  10.20.32.0/22  PGA on · NO default route to the internet
  Cloud NAT     reserved static IPs, logging on, app subnet only
  PSC endpoint  → Cloud SQL, in the data subnet
  firewall      default deny egress; allow by network tag, one reason per rule
  Cloud DNS     private zone for internal names
  flow logs     on, sampled, exported to BigQuery

Global external ALB
  Google-managed certificate
  Cloud Armor: preconfigured WAF rules in preview, plus a rate-based rule
  health check + the 35.191.0.0/16 and 130.211.0.0/22 firewall rule
  backends in two regions, RATE balancing with a measured max

The three things most likely to be wrong, and a reviewer should check them first: the health-check firewall rule (its absence makes everything 502), Private Google Access on both subnets (its absence makes private VMs hang), and whether the data subnet genuinely has no egress route rather than merely no NAT.


Topic 4: Phase 3 — Workload, Data and Keys

Workload (choose one and justify it)
  Cloud Run   ingress=internal-and-cloud-load-balancing, min-instances≥1,
              deployed by digest, secrets from Secret Manager
  or GKE      regional cluster, Workload Identity on, private nodes,
              release channel + maintenance window set

Data
  Cloud SQL   REGIONAL availability, private IP only, PITR on,
              CMEK, deletion protection, IAM database auth
  GCS         uniform bucket-level access, public access prevention,
              CMEK, lifecycle rule with a written retention reason
  Secret Manager  per-secret IAM, user-managed replication in-region

Keys
  key ring in the same region as the data
  90-day rotation, destruction delay left long
  service agents granted encrypterDecrypter — and nobody else

Prove the identity story rather than assuming it: there should be no service-account key file anywhere, the workload should reach the database through IAM authentication, and every secret access should appear in the audit log with the workload’s own identity.


Topic 5: Phase 4 — The Seven Drills

Each drill: establish the state, run it, verify with a command, record the time.

1. Kill a zone. Delete the instances in one zone of the regional MIG, or isolate it with a firewall rule. Expect: the service stays up, the load balancer stops routing there, capacity is replaced. Watch for: a zonal disk, a zonal cluster, or a Cloud SQL instance that turned out to be ZONAL.

2. Violate an org policy. Try to create a VM with an external IP. Expect: a constraint violation naming the policy. Verify: the message says constraints/…, not PERMISSION_DENIED — and you can explain the difference.

3. Restore the database with PITR. To five minutes before a deliberate bad UPDATE. Expect: a new instance with a new connection name. Record: the time from decision to a served query.

4. Rotate a CMEK key version. Create a new version, then disable the old one and observe what breaks and what does not. Expect: existing data still readable (it uses the old version) until you disable it — at which point it is not. Record: what recovered when you re-enabled.

5. Prove the data tier cannot egress. From a data-subnet host via IAP, attempt an outbound connection. Expect: a timeout. Verify: the flow log entry showing no path. This is the artifact that answers an auditor.

6. Break a health check. Make the backend return 500 on /healthz. Expect: backends go UNHEALTHY and the LB returns 502 with failed_to_pick_backend. Record: how long from first failure to first 502, and how long to recover.

7. Promote a cross-region replica. Expect: a manual promotion that permanently breaks replication, plus an application that needs repointing. Record: the full elapsed time, and write down what you would have to change to halve it.


Topic 6: What to Produce, and How It Is Judged

Four artifacts:

  1. The Terraform repository — modules with real interfaces, a GCS backend with state locking, no hardcoded project IDs, and a clean plan against the deployed state.
  2. An architecture note — one page, a decision and a rejected alternative for each significant choice. “Shared VPC rather than peering, because peering is not transitive and we will have five service projects within a year.”
  3. A runbook — the seven drills with commands, verification and measured timings, written for someone who has not seen the system.
  4. A cost breakdown — monthly cost by line item with the three largest identified. If you cannot explain the load balancer forwarding-rule charge or the inter-zone egress, that is the gap to close.

The questions you should be able to answer without notes:

  • Which single resource, if deleted, causes the largest outage?
  • What happens if the CMEK key is disabled? If it is destroyed?
  • Which of your resources are zonal, and what fails with each zone?
  • Can a service project team change the network? Should they be able to?
  • Where would a PERMISSION_DENIED come from in this design — IAM, org policy, or VPC-SC?
  • What in this build breaks first if traffic increases tenfold?

Common mistake: building the environment, screenshotting the architecture, and never breaking it. Drills 1, 4, 5 and 6 all fail on landing zones that look perfect in Terraform — the zonal disk under a regional workload, the CMEK key in the wrong region, the data subnet that turned out to have a NAT route, and the missing health-check firewall rule. The build proves it exists; the drills prove somebody can operate it.