Everything in this module, assembled once, in code, and then deliberately broken. The build is the easy half; the drills are what turn it into evidence you can show someone.
Topic 1: The Target
The requirements, stated as constraints rather than as a shopping list:
1. Survives the loss of any one availability zone, with evidence.
2. No inbound SSH from anywhere, to anything, ever.
3. No long-lived AWS credentials on any host or in any repository.
4. The database is unreachable from the internet in both directions.
5. Object storage traffic never traverses a NAT gateway.
6. Every resource is tagged well enough to attribute its cost.
7. `terraform destroy` leaves nothing behind — no orphaned volumes,
snapshots, log groups, or Elastic IPs.
Constraint 7 is not housekeeping. An environment that cannot be destroyed cleanly cannot be rebuilt reliably, and rebuild-ability is what constraint 1 depends on.
Topic 2: Phase 1 — The Network
VPC 10.20.0.0/16, three AZs
public-{a,b,c} /20 ALB nodes, NAT gateways
private-app-{a,b,c} /20 sized for pods, not instances
private-data-{a,b,c} /22 RDS only
igw one, VPC-level
nat-{a,b,c} one per AZ, each private route
table pointing at its own
s3 gateway endpoint on every private route table
interface endpoints ssm, ssmmessages, ec2messages,
logs, secretsmanager, kms
flow logs to CloudWatch Logs, 14-day retention
Decisions to make explicitly and write down:
- Why the data subnets have no
0.0.0.0/0route at all, and what that prevents that a security group does not. - Why a NAT gateway per AZ rather than one shared — both the availability reason and the cross-AZ data charge.
- Which interface endpoints you chose and what each replaces. Every endpoint you add has an hourly cost per AZ; every one you skip is either a NAT charge or a broken feature.
- Your subnet sizing arithmetic, including the usable-address count after AWS’s five reserved addresses.
Verify before moving on:
aws ec2 describe-route-tables \
--filters Name=vpc-id,Values=vpc-0abc \
--query 'RouteTables[].{id:RouteTableId,routes:Routes[].[DestinationCidrBlock,GatewayId,NatGatewayId]}'
aws ec2 describe-subnets --filters Name=vpc-id,Values=vpc-0abc \
--query 'Subnets[].[Tags[?Key==`Name`].Value|[0],CidrBlock,AvailabilityZone,AvailableIpAddressCount]' \
--output table
Topic 3: Phase 2 — Compute and the Load Balancer
Launch template
├─ AMI from a pipeline, referenced by SSM parameter, not hardcoded
├─ Instance profile: SSM core + the app's own least-privilege policy
├─ Metadata options: HttpTokens=required, hop limit 1
├─ EBS: gp3, encrypted with a customer-managed key
├─ Tags propagated to instances AND volumes
└─ User data: last-mile config from Parameter Store only
ASG
├─ min 3 / desired 3 / max 9, across all three private-app subnets
├─ health check type: ELB, grace period > measured boot-to-service
├─ target tracking on ALBRequestCountPerTarget
└─ instance refresh with MinHealthyPercentage 90 and a checkpoint
ALB
├─ public subnets, all three AZs
├─ HTTPS listener with an ACM certificate; HTTP redirects to HTTPS
├─ target group with /healthz, and a deregistration delay you chose
└─ access logs to S3
The two settings most likely to be wrong, and the ones a reviewer should check first: health check type on the ASG (EC2 is the default and it is the wrong answer), and the grace period (guessed from boot time rather than measured to first passing health check).
Measure, do not estimate: launch one instance from the template and time RunInstances to first healthy target. That number sets the grace period and is also your scale-out latency.
Topic 4: Phase 3 — Data and Secrets
RDS
├─ Multi-AZ, in the data subnets, PubliclyAccessible=false
├─ Encrypted at creation with a customer-managed key
├─ Security group: inbound 5432 from sg-app only — not a CIDR
├─ Automated backups, retention chosen against a written RPO
├─ Deletion protection on
├─ Performance Insights and enhanced monitoring on
└─ force_ssl enforced in the parameter group
Secrets Manager
├─ The database credential, with rotation enabled
├─ Retrieved by the app through its instance profile — never in user data
└─ Reached through the secretsmanager interface endpoint
S3
├─ Block Public Access on, ACLs disabled (BucketOwnerEnforced)
├─ Versioning on, with a lifecycle rule expiring noncurrent versions
├─ A bucket policy denying non-TLS and denying outside the org
└─ Reached through the gateway endpoint, with an endpoint policy
restricting it to your buckets
The rotation test is the one that finds real bugs. Force a rotation while traffic is flowing and watch what the application does. An application that fetched the secret once at startup fails at the next reconnect, hours later, with nothing to correlate against — which is exactly why this belongs in the build rather than in a later incident.
Topic 5: Phase 4 — The Five Drills
This is the part that makes the project worth doing. Each drill has an expected result and a number you must record.
Drill 1 — Kill an availability zone. Use Fault Injection Service, or remove one AZ’s subnets from the ALB and stop its instances. Expect: the service stays up on two AZs, the ALB stops routing to the lost one, the ASG relaunches in the survivors. Record: the error rate during the transition and its duration. Watch for: a single shared NAT gateway, uneven target counts with cross-zone off, and quotas that prevent the remaining AZs from taking the load.
Drill 2 — Fail the application health check. Make /healthz return 500 on one instance.
Expect: the ALB removes it within (interval × unhealthy threshold), the ASG replaces it because health check type is ELB.
Record: time from first failure to a replacement instance serving traffic. That is your real self-healing time.
Drill 3 — Restore the database. Point-in-time restore to five minutes ago, into a new instance, and repoint the application. Expect: a new endpoint, not a rewind. Record: time from decision to a served request, decision time included. That is your RTO, and it is usually larger than you assumed.
Drill 4 — Rotate the secret under load. Force rotation while requests flow. Expect: no errors. Record: whether the application reconnected, and if it did not, exactly why.
Drill 5 — Prove the data tier is isolated. From a database-tier host (via Session Manager), attempt an outbound connection to the internet. Expect: a timeout, with a flow log entry showing no path. Record: the flow log line. It is the artifact that answers an auditor.
Topic 6: What to Produce, and How It Is Judged
Four artifacts:
- The Terraform repository — modules with real interfaces, a remote backend with locking, no hardcoded AMI IDs or account numbers, and a clean
planagainst the deployed state. - An architecture note — one page, with a decision and a rejected alternative for each significant choice. “NAT per AZ, rejected shared NAT because an AZ failure removes egress for the other two and cross-AZ NAT traffic bills twice.”
- A runbook — the five drills, with commands and measured timings, written for someone who has not seen the system.
- A cost breakdown — the monthly cost by line item, with the three largest identified. If NAT gateway data processing is not in your top three, explain what you did to keep it out.
The questions you should be able to answer without notes, which are also the questions an interviewer will ask:
- Which single resource, if deleted, causes the largest outage? (Usually a route table or the ALB — and it should not be a NAT gateway.)
- What happens if the KMS key is scheduled for deletion?
- Where does a request’s client IP appear in your logs, and what preserved it?
- How would this change if the database had to survive losing the whole region?
- Which of these resources are zonal, and what fails with each AZ?
- What in this build would break first if traffic increased tenfold?
Common mistake: building the environment, taking a screenshot of the architecture, and never breaking it. Every part of this design is defensible on paper and only some of it works — the health check type, the grace period, the secret rotation and the multi-AZ NAT are all things that look correct in Terraform and fail in a drill. The drills are the deliverable; the infrastructure is just what they run against.