☁️
Active Module Path

AWS Operations

Run AWS the way it actually fails β€” the blast-radius model, IAM evaluation order, VPC routing, EC2 and EBS ceilings, RDS failover, EKS identity and capacity, and a landing zone you break on purpose.

20 lessons 445 min syllabus

Stage 1 β€” Account & Identity

4 lessons
Lesson 1 β€’ ⏱️ 18m
Regions, AZs and the Blast Radius Model

Why every AWS design decision reduces to one question β€” what dies with this resource β€” and how the global, regional and zonal scopes answer it.

βœ“
Lesson 2 β€’ ⏱️ 22m
IAM: How a Request Is Actually Authorized

The evaluation order AWS runs on every API call, why an explicit Deny cannot be argued with, and how to read an AccessDenied message instead of guessing at it.

βœ“
Lesson 3 β€’ ⏱️ 20m
Roles, STS and the Credential Chain

Where the credentials in your terminal actually came from, why a role is not a user with extra steps, and how IMDSv2 turned an SSRF bug into a non-event.

βœ“
Lesson 4 β€’ ⏱️ 20m
Organizations, SCPs and the Account Boundary

Why the AWS account is the real isolation boundary, what a service control policy can and cannot do, and how to lay out accounts so a mistake stays inside one of them.

βœ“

Stage 2 β€” The Network

4 lessons
Lesson 5 β€’ ⏱️ 20m
VPC Addressing and Subnet Design

Sizing a VPC you will not have to rebuild: CIDR arithmetic, the five addresses AWS takes from every subnet, and why pod density β€” not instance count β€” decides your subnet size.

βœ“
Lesson 6 β€’ ⏱️ 20m
Routing: Internet Gateway, NAT and Egress

What actually makes a subnet public, why a NAT gateway is a one-way door with an hourly bill, and how to read a route table the way the VPC router reads it.

βœ“
Lesson 7 β€’ ⏱️ 18m
Security Groups and Network ACLs

Stateful versus stateless, why a NACL breaks return traffic you never thought about, and the security group referencing pattern that survives every subnet change.

βœ“
Lesson 8 β€’ ⏱️ 22m
Endpoints, Peering and Transit Gateway

Reaching AWS services and other VPCs without touching the internet β€” gateway versus interface endpoints, why peering does not scale past three VPCs, and what a Transit Gateway route table is really for.

βœ“

Stage 3 β€” Compute

3 lessons
Lesson 9 β€’ ⏱️ 22m
EC2 Instances: Lifecycle, Families and Placement

Reading an instance type name, what each state actually bills, why stop-start moves your host, and the bootstrapping choices that decide how fast a replacement enters service.

βœ“
Lesson 10 β€’ ⏱️ 20m
Block Storage: EBS, Instance Store and Snapshots

The two ceilings on EBS performance, why gp3 replaced gp2 for almost everything, and what a crash-consistent snapshot actually promises about your database.

βœ“
Lesson 11 β€’ ⏱️ 22m
Load Balancing and Auto Scaling

Choosing between ALB and NLB on their real differences, the two independent health checks that decide whether a broken fleet ever heals, and scaling policies that react before users do.

βœ“

Stage 4 β€” State & Data

3 lessons
Lesson 12 β€’ ⏱️ 22m
S3: Durability, Classes and Access Control

What eleven nines actually promises, which storage class is a trap for small objects, and the four independent mechanisms that decide whether a request to a bucket succeeds.

βœ“
Lesson 13 β€’ ⏱️ 22m
RDS and Aurora Operations

Why Multi-AZ adds no read capacity, what point-in-time recovery really restores, and the maintenance and failover behaviours that decide whether your application notices.

βœ“
Lesson 14 β€’ ⏱️ 20m
KMS, Encryption and Secrets

Envelope encryption in one diagram, why the key policy outranks IAM, and choosing between Secrets Manager and Parameter Store on something other than price.

βœ“

Stage 5 β€” Kubernetes on AWS

2 lessons
Lesson 15 β€’ ⏱️ 24m
EKS Cluster Anatomy and Node Capacity

What AWS runs and what you still own, the pod-per-node IP arithmetic that decides your subnet sizing, and the endpoint-access setting that can lock everyone out of a healthy cluster.

βœ“
Lesson 16 β€’ ⏱️ 22m
IRSA, Pod Identity and Cluster Access

How a Kubernetes ServiceAccount becomes an AWS identity, the trust policy condition that silently grants too much, and why aws-auth was replaced by access entries.

βœ“

Stage 6 β€” Operations

4 lessons
Lesson 17 β€’ ⏱️ 22m
Observability: CloudWatch, CloudTrail and Config

Three services that answer three different questions, the metric-filter and composite-alarm patterns worth building once, and why alarming on utilisation produces pages nobody acts on.

βœ“
Lesson 18 β€’ ⏱️ 22m
Resilience and DR: RTO, RPO and the Four Strategies

Turning an availability conversation into two numbers, choosing the cheapest strategy that meets them, and why an untested plan has an unknown recovery time rather than a slow one.

βœ“
Lesson 19 β€’ ⏱️ 22m
The AWS Failure Playbook

A fixed six-step sweep for any AWS incident, the four error classes and what each actually means, and the quota and throttling failures that present as random.

βœ“
Lesson 20 β€’ ⏱️ 45m
Project: A Production Three-Tier Landing Zone

Build the whole module as one environment β€” a three-AZ VPC, a private application tier with no bastion, an encrypted Multi-AZ database β€” then run five drills that prove it recovers.

βœ“

πŸ—ΊοΈ Beginner β†’ Expert Roadmap

6 stages with prerequisites and a concrete mastery check at each.

→

🎯 What You'll Learn

  • β€’ Classify any resource as global, regional or zonal, and name what dies with it.
  • β€’ Explain an AccessDenied by the evaluation stage that produced it, instead of adding permissions until it works.
  • β€’ Say where the credentials in any shell, instance or pod came from β€” and prove it with one command.
  • β€’ Write SCPs as guardrails that an account administrator cannot argue around.
  • β€’ Size a VPC for pod density rather than instance count, and never rebuild it.
  • β€’ Read a route table the way the VPC router reads it, and know what makes a subnet public.
  • β€’ Choose between security groups and NACLs on statefulness, and reference groups instead of CIDRs.
  • β€’ Cut NAT cost and a dependency at once with the endpoints that should already exist.
  • β€’ Diagnose whether the volume or the instance is the EBS ceiling, with a measurement.
  • β€’ Configure an ASG that actually replaces an instance the load balancer has given up on.
  • β€’ Set S3 controls that survive a later policy mistake, and spot the lifecycle rule that raises the bill.
  • β€’ Distinguish Multi-AZ from read replicas, and time a failover your application really survives.
  • β€’ Explain why a key policy outranks IAM, and keep an application alive through a secret rotation.
  • β€’ Calculate EKS pod capacity from ENI limits, and write an IRSA trust policy that grants exactly one service account.
  • β€’ Alarm on user-facing symptoms rather than utilisation, and run the CloudWatch β†’ CloudTrail β†’ Config sequence in order.
  • β€’ Turn an availability requirement into an RTO and an RPO, then pick the cheapest strategy that meets both.
  • β€’ Run a fixed six-step triage sweep instead of debugging by intuition.
  • β€’ Ship a three-tier landing zone in Terraform and prove it with five drills, not a screenshot.

πŸ›‘οΈ Best Practices in Production

The short version of this path. Every lesson also ends with the specific mistake it exists to prevent.

Do this
  • βœ“ Classify every resource as global, regional or zonal, and know what dies with each.
  • βœ“ One NAT gateway per AZ, with each private route table pointing at its own.
  • βœ“ Reference security groups rather than CIDRs, so rules survive subnet changes.
  • βœ“ Add the free S3 and DynamoDB gateway endpoints, and interface endpoints for SSM.
  • βœ“ Enforce IMDSv2 with hop limit 1, and use roles β€” never long-lived access keys.
  • βœ“ Write organisation guardrails as SCP Denies, exempting global services from region restrictions.
  • βœ“ Alarm on user-facing symptoms and error rates, not on CPU utilisation.
  • βœ“ Turn an availability requirement into a written RTO and RPO, then test the restore.
Avoid this
  • βœ— Sizing VPC subnets by instance count when EKS gives every pod a real IP.
  • βœ— Debugging AccessDenied by adding permissions until it works, instead of reading which policy denied it.
  • βœ— Encrypting with an AWS-managed KMS key, then needing to share the snapshot cross-account.
  • βœ— An ASG left on the default EC2 health check, so it never replaces a target the ALB gave up on.
  • βœ— Attaching a region-deny SCP to accounts with existing resources in those regions β€” they become unmanageable.
  • βœ— Provisioning IOPS the instance size cannot carry.

πŸ’Ό Interview Readiness

Once you reach the end of this path, test your engineering knowledge against real questions asked by top technical teams:

A multi-region system's architecture team wants to justify why they've deliberately built in the assumption that any given region can fail at any time, rather than treating regional failure as a rare edge case. How would you frame this design philosophy to a skeptical stakeholder who sees it as over-engineering?

Career & Interview Strategy

Explain the "home-region ownership" pattern and how it prevents a double-withdrawal scenario in a multi-region banking system.

Career & Interview Strategy

A client asks for a "rollback plan" before you delete/filter historical log data. What do you need to clarify with them before proceeding?

Cloud Cost Optimization

Why can you NOT specify an IAM instance profile in the Launch Template when creating EKS node groups via CLI?

Cloud Cost Optimization

How do you implement FinOps for a multi-account AWS organization with 20 accounts?

Cloud Cost Optimization

A pod is stuck in Pending with no node or IP assigned. Should you check CNI logs first?

Kubernetes

SSH times out connecting to an EC2 instance. Ping also fails. But nc -zv <ip> 22 succeeds. What does this tell you?

Linux & Networking

Four ways to access an EC2 instance when SSH is broken. List them in order of preference and explain the tradeoff.

Linux & Networking

A client wants a 50/50 multi-cloud split "for cost savings." How would you push back or reframe this conversation?

Cloud Cost Optimization

What is an EC2 Launch Template and why does it matter for ASG migrations?

Cloud Cost Optimization

How do you migrate 170 EC2 instances from Intel to AMD with zero unplanned downtime?

Cloud Cost Optimization

When is it appropriate to stop EC2 instances to save cost, and when is it not?

Cloud Cost Optimization

A production incident shows CPU, RAM, and disk all reporting as "healthy" in monitoring, yet a specific service is clearly struggling. What's a diagnostic angle that basic resource monitoring might miss?

Structured Debugging

What is a cluster autoscaler, and why does a Kubernetes cluster need one?

Cloud Cost Optimization

What is an S3 (or GCS) lifecycle policy, and what problem does it solve?

Cloud Cost Optimization

A client asks you to confirm that reducing a log bucket's retention period from 30 to 7 days won't affect their compliance posture. How would you verify and communicate this?

Cloud Cost Optimization

Why does this system avoid synchronous cross-region database writes?

Career & Interview Strategy

How does the CAP theorem relate to the multi-region architecture decisions described in this guide?

Career & Interview Strategy

A prospective client is based in the European Union and has strict data-residency requirements. You're evaluating CastAI (US-only SaaS) versus Karpenter (open-source, in-cluster) for their EKS autoscaling needs. Walk through your decision process.

Kubernetes Autoscaling

How does Karpenter differ from a standard EKS managed node group with Cluster Autoscaler?

Cloud Cost Optimization

Why does EKS node group migration require ~2–3 minutes of downtime, while ASG migration has zero downtime?

Cloud Cost Optimization

Design a node pool strategy for an EKS cluster that needs to support both a real-time customer-facing calculation service and a large nightly Spark analytics job, using the principles discussed in this guide.

Kubernetes Autoscaling

What are the three EKS security scanning tools covered and what does each focus on?

Cloud Cost Optimization

How does Karpenter know which subnets and security groups to use when provisioning a new node?

Kubernetes Autoscaling

A client refuses to let you make application-level logging changes, but wants logging costs reduced. What are your levers, and what are the trade-offs of each?

Cloud Cost Optimization

A client insists on continuing to run both a self-hosted Prometheus/Grafana stack and a cloud-native monitoring tool in parallel, citing team preference. How would you approach cost optimization given this constraint, rather than pushing for consolidation?

Cloud Cost Optimization

Why should multi-region deployments or patching operations always be sequential rather than parallel?

CI/CD & Automation

Why is rollback done in parallel across all regions while forward deployment is sequential?

CI/CD & Automation

A service status check that normally responds instantly is now taking 8 seconds to respond. What would you conclude, and what would you do next?

Structured Debugging

A junior engineer on your team says a system issue "just fixed itself" during a live troubleshooting session, and they can no longer reproduce a bug you were actively diagnosing together. What questions would you ask before concluding the issue is actually resolved?

Structured Debugging

A client has a $98K/month AWS bill. Walk me through how you'd approach reducing it.

Cloud Cost Optimization

What tags should every AWS resource have, and why?

Cloud Cost Optimization