Project: A Production Three-Tier Landing Zone

Build the whole module as one environment — a three-AZ VPC, a private application tier with no bastion, an encrypted Multi-AZ database — then run five drills that prove it recovers.

advanced 45 min lesson hands-on task included

Everything in this module, assembled once, in code, and then deliberately broken. The build is the easy half; the drills are what turn it into evidence you can show someone.


Topic 1: The Target

PROJECT TARGET — EVERY BOX IS SOMETHING YOU MUST BE ABLE TO DEFEND users VPC 10.20.0.0/16 · 3 AZs · flow logs on PUBLIC SUBNETS ALB + ACM cert NAT GW × AZ no EC2 lives here PRIVATE APP SUBNETS ASG · launch template · IMDSv2 instance profile, no keys on disk health check type: ELB interface endpoints ssm · ssmmessages · ec2messages logs · secretsmanager · ecr.* PRIVATE DATA SUBNETS RDS Multi-AZ, encrypted Secrets Manager rotation S3 gateway endpoint — the route that stops paying NAT for object traffic NO BASTION Access is ssm start-session No port 22 open anywhere, to anyone. Audited in CT. PROVE IT □ kill an AZ□ fail the /healthz□ restore the DB□ rotate the secret□ read the flow log Every AWS lesson in this module appears somewhere in this diagram. That is the point of it. Build it with Terraform, not the console — the console version cannot be reviewed, and cannot be destroyed cleanly.
Every box here corresponds to a lesson in this module. If you cannot say which lesson explains a box and why the alternative was rejected, that is the lesson to re-read before building it.

The requirements, stated as constraints rather than as a shopping list:

1. Survives the loss of any one availability zone, with evidence.
2. No inbound SSH from anywhere, to anything, ever.
3. No long-lived AWS credentials on any host or in any repository.
4. The database is unreachable from the internet in both directions.
5. Object storage traffic never traverses a NAT gateway.
6. Every resource is tagged well enough to attribute its cost.
7. `terraform destroy` leaves nothing behind — no orphaned volumes,
   snapshots, log groups, or Elastic IPs.

Constraint 7 is not housekeeping. An environment that cannot be destroyed cleanly cannot be rebuilt reliably, and rebuild-ability is what constraint 1 depends on.


Topic 2: Phase 1 — The Network

VPC 10.20.0.0/16, three AZs
  public-{a,b,c}        /20   ALB nodes, NAT gateways
  private-app-{a,b,c}   /20   sized for pods, not instances
  private-data-{a,b,c}  /22   RDS only

  igw                          one, VPC-level
  nat-{a,b,c}                  one per AZ, each private route
                               table pointing at its own
  s3 gateway endpoint          on every private route table
  interface endpoints          ssm, ssmmessages, ec2messages,
                               logs, secretsmanager, kms
  flow logs                    to CloudWatch Logs, 14-day retention

Decisions to make explicitly and write down:

  • Why the data subnets have no 0.0.0.0/0 route at all, and what that prevents that a security group does not.
  • Why a NAT gateway per AZ rather than one shared — both the availability reason and the cross-AZ data charge.
  • Which interface endpoints you chose and what each replaces. Every endpoint you add has an hourly cost per AZ; every one you skip is either a NAT charge or a broken feature.
  • Your subnet sizing arithmetic, including the usable-address count after AWS’s five reserved addresses.

Verify before moving on:

aws ec2 describe-route-tables \
  --filters Name=vpc-id,Values=vpc-0abc \
  --query 'RouteTables[].{id:RouteTableId,routes:Routes[].[DestinationCidrBlock,GatewayId,NatGatewayId]}'
aws ec2 describe-subnets --filters Name=vpc-id,Values=vpc-0abc \
  --query 'Subnets[].[Tags[?Key==`Name`].Value|[0],CidrBlock,AvailabilityZone,AvailableIpAddressCount]' \
  --output table

Topic 3: Phase 2 — Compute and the Load Balancer

Launch template
  ├─ AMI from a pipeline, referenced by SSM parameter, not hardcoded
  ├─ Instance profile: SSM core + the app's own least-privilege policy
  ├─ Metadata options: HttpTokens=required, hop limit 1
  ├─ EBS: gp3, encrypted with a customer-managed key
  ├─ Tags propagated to instances AND volumes
  └─ User data: last-mile config from Parameter Store only

ASG
  ├─ min 3 / desired 3 / max 9, across all three private-app subnets
  ├─ health check type: ELB, grace period > measured boot-to-service
  ├─ target tracking on ALBRequestCountPerTarget
  └─ instance refresh with MinHealthyPercentage 90 and a checkpoint

ALB
  ├─ public subnets, all three AZs
  ├─ HTTPS listener with an ACM certificate; HTTP redirects to HTTPS
  ├─ target group with /healthz, and a deregistration delay you chose
  └─ access logs to S3

The two settings most likely to be wrong, and the ones a reviewer should check first: health check type on the ASG (EC2 is the default and it is the wrong answer), and the grace period (guessed from boot time rather than measured to first passing health check).

Measure, do not estimate: launch one instance from the template and time RunInstances to first healthy target. That number sets the grace period and is also your scale-out latency.


Topic 4: Phase 3 — Data and Secrets

RDS
  ├─ Multi-AZ, in the data subnets, PubliclyAccessible=false
  ├─ Encrypted at creation with a customer-managed key
  ├─ Security group: inbound 5432 from sg-app only — not a CIDR
  ├─ Automated backups, retention chosen against a written RPO
  ├─ Deletion protection on
  ├─ Performance Insights and enhanced monitoring on
  └─ force_ssl enforced in the parameter group

Secrets Manager
  ├─ The database credential, with rotation enabled
  ├─ Retrieved by the app through its instance profile — never in user data
  └─ Reached through the secretsmanager interface endpoint

S3
  ├─ Block Public Access on, ACLs disabled (BucketOwnerEnforced)
  ├─ Versioning on, with a lifecycle rule expiring noncurrent versions
  ├─ A bucket policy denying non-TLS and denying outside the org
  └─ Reached through the gateway endpoint, with an endpoint policy
     restricting it to your buckets

The rotation test is the one that finds real bugs. Force a rotation while traffic is flowing and watch what the application does. An application that fetched the secret once at startup fails at the next reconnect, hours later, with nothing to correlate against — which is exactly why this belongs in the build rather than in a later incident.


Topic 5: Phase 4 — The Five Drills

This is the part that makes the project worth doing. Each drill has an expected result and a number you must record.

Drill 1 — Kill an availability zone. Use Fault Injection Service, or remove one AZ’s subnets from the ALB and stop its instances. Expect: the service stays up on two AZs, the ALB stops routing to the lost one, the ASG relaunches in the survivors. Record: the error rate during the transition and its duration. Watch for: a single shared NAT gateway, uneven target counts with cross-zone off, and quotas that prevent the remaining AZs from taking the load.

Drill 2 — Fail the application health check. Make /healthz return 500 on one instance. Expect: the ALB removes it within (interval × unhealthy threshold), the ASG replaces it because health check type is ELB. Record: time from first failure to a replacement instance serving traffic. That is your real self-healing time.

Drill 3 — Restore the database. Point-in-time restore to five minutes ago, into a new instance, and repoint the application. Expect: a new endpoint, not a rewind. Record: time from decision to a served request, decision time included. That is your RTO, and it is usually larger than you assumed.

Drill 4 — Rotate the secret under load. Force rotation while requests flow. Expect: no errors. Record: whether the application reconnected, and if it did not, exactly why.

Drill 5 — Prove the data tier is isolated. From a database-tier host (via Session Manager), attempt an outbound connection to the internet. Expect: a timeout, with a flow log entry showing no path. Record: the flow log line. It is the artifact that answers an auditor.


Topic 6: What to Produce, and How It Is Judged

Four artifacts:

  1. The Terraform repository — modules with real interfaces, a remote backend with locking, no hardcoded AMI IDs or account numbers, and a clean plan against the deployed state.
  2. An architecture note — one page, with a decision and a rejected alternative for each significant choice. “NAT per AZ, rejected shared NAT because an AZ failure removes egress for the other two and cross-AZ NAT traffic bills twice.”
  3. A runbook — the five drills, with commands and measured timings, written for someone who has not seen the system.
  4. A cost breakdown — the monthly cost by line item, with the three largest identified. If NAT gateway data processing is not in your top three, explain what you did to keep it out.

The questions you should be able to answer without notes, which are also the questions an interviewer will ask:

  • Which single resource, if deleted, causes the largest outage? (Usually a route table or the ALB — and it should not be a NAT gateway.)
  • What happens if the KMS key is scheduled for deletion?
  • Where does a request’s client IP appear in your logs, and what preserved it?
  • How would this change if the database had to survive losing the whole region?
  • Which of these resources are zonal, and what fails with each AZ?
  • What in this build would break first if traffic increased tenfold?

Common mistake: building the environment, taking a screenshot of the architecture, and never breaking it. Every part of this design is defensible on paper and only some of it works — the health check type, the grace period, the secret rotation and the multi-AZ NAT are all things that look correct in Terraform and fail in a drill. The drills are the deliverable; the infrastructure is just what they run against.