Identity → Operations

AWS Operations Roadmap

Six stages, each gated by a capability rather than a lesson count. This is an operations path, not a certification path — every stage ends with something you can either do or cannot, and the last one ends with an environment you break on purpose.

20
Lessons
1
Projects
7h
Total
6
Stages

Dependencies that are not optional

Most lessons can be taken in any order within a stage. These four cannot.

  • IAM evaluation order before everything else Half of all AWS failures are permission failures wearing a different costume. You cannot debug the rest until you can read a denial.
  • Subnet sizing before EKS capacity The VPC CNI gives every pod a subnet IP. A subnet sized for nodes runs out during a deployment, and the only quick fix requires recycling every node.
  • Blast radius before any DR conversation RTO and RPO are meaningless until you can say which resources are zonal. Most "we need multi-region" is an unresolved single-AZ dependency.
  • Health checks before auto scaling An ASG on the default EC2 health check never replaces an instance the load balancer has already removed. The fleet stays "healthy" and serves nothing.

Stage 1 — Account & Identity

~1.5h

Know what fails together, and who is allowed to touch it

4 lessons
Prerequisite

An AWS account you can create resources in. Basic Linux, and comfort with a terminal.

You will be able to
  • ▸Classify any resource as global, regional or zonal and name what dies with it
  • ▸Walk a request through the five stages of IAM evaluation from memory
  • ▸Say where the credentials in any shell, instance or pod came from
  • ▸Enforce IMDSv2 and explain what it closed
  • ▸Write an SCP guardrail an account administrator cannot argue around
Mastery check — can you do this?

Handed an AccessDenied you have never seen, you name the evaluation stage that produced it and open the right policy first time — and you can prove the answer with the policy simulator before changing anything.

Stage 2 — The Network

~1.5h

Follow the packet; stop guessing at the console

4 lessons
Prerequisite

Stage 1. Enough CIDR arithmetic to halve a block in your head.

You will be able to
  • ▸Size a VPC for pod density rather than instance count
  • ▸Read a route table the way the VPC router reads it, longest prefix first
  • ▸Explain what actually makes a subnet public, and why a NAT gateway is zonal
  • ▸Choose security groups over NACLs deliberately, and reference groups instead of CIDRs
  • ▸Reach AWS services privately, and know when peering has stopped scaling
Mastery check — can you do this?

Given "the private instance cannot reach the internet", you find the cause in four commands without opening the console — and your data subnets have no default route at all, on purpose.

Stage 3 — Compute

~1h

Capacity that heals itself, and storage that is not silently capped

3 lessons
Prerequisite

Stage 2 — a load balancer and an ASG are network objects before they are compute ones.

You will be able to
  • ▸Read an instance type name and predict its network and EBS ceilings
  • ▸Measure boot-to-service time and set a grace period from it rather than from a guess
  • ▸Prove whether the volume or the instance limits your I/O
  • ▸Configure an ASG that replaces an instance the load balancer has given up on
  • ▸Roll a fleet with instance refresh and stop it safely mid-rollout
Mastery check — can you do this?

You break the application health check on every instance and the fleet heals itself without intervention — and you can state, in seconds, how long that took and which setting decided it.

Stage 4 — State & Data

~1h

The part you cannot roll back

3 lessons
Prerequisite

Stage 1 for the policy work; stage 2 for where a database is allowed to live.

You will be able to
  • ▸Set S3 controls that survive a later policy mistake
  • ▸Spot the lifecycle rule that increases the bill instead of reducing it
  • ▸Distinguish Multi-AZ from read replicas and use each for its actual purpose
  • ▸Restore to a point in time and know that it creates a new endpoint
  • ▸Explain why a key policy outranks IAM, and keep an app alive through a rotation
Mastery check — can you do this?

You force a Multi-AZ failover and a secret rotation on a live application, and neither produces an error your code did not already handle.

Stage 5 — Kubernetes on AWS

~1h

Kubernetes, with AWS underneath it

2 lessons
Prerequisite

Stages 1–4, plus the Kubernetes path through workloads and scheduling.

This stage assumes Kubernetes itself. Everything here is what AWS adds and takes away — the control-plane split, the IP budget, and identity.

You will be able to
  • ▸Calculate pods per node from ENI limits and verify against the node
  • ▸Recognise IP exhaustion from its error message rather than blaming the scheduler
  • ▸Choose between managed node groups, Karpenter and Fargate on their real trade-offs
  • ▸Write an IRSA trust policy that grants exactly one service account
  • ▸Manage cluster access with access entries instead of editing a ConfigMap by hand
Mastery check — can you do this?

You give one pod access to one bucket with no key anywhere, block the node role from that bucket, and prove that a pod without the annotation is denied.

Stage 6 — Operations

~2h

Be the person called when the account is the problem

4 lessons
Prerequisite

Stages 1–5. The project assumes every one of them.

You will be able to
  • ▸Alarm on user-facing symptoms rather than utilisation
  • ▸Run the CloudWatch → Insights → CloudTrail → Config sequence in that order
  • ▸Turn an availability requirement into an RTO and an RPO, then price both
  • ▸Recognise throttling, quota and burst-credit failures from their signatures
  • ▸Run a fixed six-step sweep instead of debugging by intuition
  • ▸Ship a landing zone in Terraform and prove it with drills
Mastery check — can you do this?

Project — you build the three-tier landing zone, remove an availability zone on purpose, and the service stays up while you read the measured numbers off your own runbook.

Scope, and what changes underneath it

AWS ships features continuously, so any document about it is partly out of date the week it is written. This path is built around the parts that have not moved in a decade — and it says so explicitly wherever a specific default or limit is quoted.

Durable
Evaluation order, blast radius, statefulness, envelope encryption, the control-plane/data-plane split.
Moves slowly
Service quotas, instance families, storage classes, add-on versions. Check them; do not memorise them.
Deliberately excluded
Cost engineering has its own path, and Kubernetes itself has another. This one cross-links rather than duplicates.

Where a lesson quotes a limit or a default, verify it against your own account rather than trusting any document, including this one: aws service-quotas list-service-quotas and aws ec2 describe-account-attributes are the two commands that settle most arguments.