AWS Operations
Run AWS the way it actually fails β the blast-radius model, IAM evaluation order, VPC routing, EC2 and EBS ceilings, RDS failover, EKS identity and capacity, and a landing zone you break on purpose.
Stage 1 β Account & Identity
4 lessonsWhy every AWS design decision reduces to one question β what dies with this resource β and how the global, regional and zonal scopes answer it.
The evaluation order AWS runs on every API call, why an explicit Deny cannot be argued with, and how to read an AccessDenied message instead of guessing at it.
Where the credentials in your terminal actually came from, why a role is not a user with extra steps, and how IMDSv2 turned an SSRF bug into a non-event.
Why the AWS account is the real isolation boundary, what a service control policy can and cannot do, and how to lay out accounts so a mistake stays inside one of them.
Stage 2 β The Network
4 lessonsSizing a VPC you will not have to rebuild: CIDR arithmetic, the five addresses AWS takes from every subnet, and why pod density β not instance count β decides your subnet size.
What actually makes a subnet public, why a NAT gateway is a one-way door with an hourly bill, and how to read a route table the way the VPC router reads it.
Stateful versus stateless, why a NACL breaks return traffic you never thought about, and the security group referencing pattern that survives every subnet change.
Reaching AWS services and other VPCs without touching the internet β gateway versus interface endpoints, why peering does not scale past three VPCs, and what a Transit Gateway route table is really for.
Stage 3 β Compute
3 lessonsReading an instance type name, what each state actually bills, why stop-start moves your host, and the bootstrapping choices that decide how fast a replacement enters service.
The two ceilings on EBS performance, why gp3 replaced gp2 for almost everything, and what a crash-consistent snapshot actually promises about your database.
Choosing between ALB and NLB on their real differences, the two independent health checks that decide whether a broken fleet ever heals, and scaling policies that react before users do.
Stage 4 β State & Data
3 lessonsWhat eleven nines actually promises, which storage class is a trap for small objects, and the four independent mechanisms that decide whether a request to a bucket succeeds.
Why Multi-AZ adds no read capacity, what point-in-time recovery really restores, and the maintenance and failover behaviours that decide whether your application notices.
Envelope encryption in one diagram, why the key policy outranks IAM, and choosing between Secrets Manager and Parameter Store on something other than price.
Stage 5 β Kubernetes on AWS
2 lessonsWhat AWS runs and what you still own, the pod-per-node IP arithmetic that decides your subnet sizing, and the endpoint-access setting that can lock everyone out of a healthy cluster.
How a Kubernetes ServiceAccount becomes an AWS identity, the trust policy condition that silently grants too much, and why aws-auth was replaced by access entries.
Stage 6 β Operations
4 lessonsThree services that answer three different questions, the metric-filter and composite-alarm patterns worth building once, and why alarming on utilisation produces pages nobody acts on.
Turning an availability conversation into two numbers, choosing the cheapest strategy that meets them, and why an untested plan has an unknown recovery time rather than a slow one.
A fixed six-step sweep for any AWS incident, the four error classes and what each actually means, and the quota and throttling failures that present as random.
Build the whole module as one environment β a three-AZ VPC, a private application tier with no bastion, an encrypted Multi-AZ database β then run five drills that prove it recovers.
πΊοΈ Beginner β Expert Roadmap
6 stages with prerequisites and a concrete mastery check at each.
π― What You'll Learn
- β’ Classify any resource as global, regional or zonal, and name what dies with it.
- β’ Explain an AccessDenied by the evaluation stage that produced it, instead of adding permissions until it works.
- β’ Say where the credentials in any shell, instance or pod came from β and prove it with one command.
- β’ Write SCPs as guardrails that an account administrator cannot argue around.
- β’ Size a VPC for pod density rather than instance count, and never rebuild it.
- β’ Read a route table the way the VPC router reads it, and know what makes a subnet public.
- β’ Choose between security groups and NACLs on statefulness, and reference groups instead of CIDRs.
- β’ Cut NAT cost and a dependency at once with the endpoints that should already exist.
- β’ Diagnose whether the volume or the instance is the EBS ceiling, with a measurement.
- β’ Configure an ASG that actually replaces an instance the load balancer has given up on.
- β’ Set S3 controls that survive a later policy mistake, and spot the lifecycle rule that raises the bill.
- β’ Distinguish Multi-AZ from read replicas, and time a failover your application really survives.
- β’ Explain why a key policy outranks IAM, and keep an application alive through a secret rotation.
- β’ Calculate EKS pod capacity from ENI limits, and write an IRSA trust policy that grants exactly one service account.
- β’ Alarm on user-facing symptoms rather than utilisation, and run the CloudWatch β CloudTrail β Config sequence in order.
- β’ Turn an availability requirement into an RTO and an RPO, then pick the cheapest strategy that meets both.
- β’ Run a fixed six-step triage sweep instead of debugging by intuition.
- β’ Ship a three-tier landing zone in Terraform and prove it with five drills, not a screenshot.
π‘οΈ Best Practices in Production
The short version of this path. Every lesson also ends with the specific mistake it exists to prevent.
- β Classify every resource as global, regional or zonal, and know what dies with each.
- β One NAT gateway per AZ, with each private route table pointing at its own.
- β Reference security groups rather than CIDRs, so rules survive subnet changes.
- β Add the free S3 and DynamoDB gateway endpoints, and interface endpoints for SSM.
- β Enforce IMDSv2 with hop limit 1, and use roles β never long-lived access keys.
- β Write organisation guardrails as SCP Denies, exempting global services from region restrictions.
- β Alarm on user-facing symptoms and error rates, not on CPU utilisation.
- β Turn an availability requirement into a written RTO and RPO, then test the restore.
- β Sizing VPC subnets by instance count when EKS gives every pod a real IP.
- β Debugging AccessDenied by adding permissions until it works, instead of reading which policy denied it.
- β Encrypting with an AWS-managed KMS key, then needing to share the snapshot cross-account.
- β An ASG left on the default
EC2health check, so it never replaces a target the ALB gave up on. - β Attaching a region-deny SCP to accounts with existing resources in those regions β they become unmanageable.
- β Provisioning IOPS the instance size cannot carry.
πΌ Interview Readiness
Once you reach the end of this path, test your engineering knowledge against real questions asked by top technical teams:
A multi-region system's architecture team wants to justify why they've deliberately built in the assumption that any given region can fail at any time, rather than treating regional failure as a rare edge case. How would you frame this design philosophy to a skeptical stakeholder who sees it as over-engineering?
Career & Interview StrategyExplain the "home-region ownership" pattern and how it prevents a double-withdrawal scenario in a multi-region banking system.
Career & Interview StrategyA client asks for a "rollback plan" before you delete/filter historical log data. What do you need to clarify with them before proceeding?
Cloud Cost OptimizationWhy can you NOT specify an IAM instance profile in the Launch Template when creating EKS node groups via CLI?
Cloud Cost OptimizationHow do you implement FinOps for a multi-account AWS organization with 20 accounts?
Cloud Cost OptimizationA pod is stuck in Pending with no node or IP assigned. Should you check CNI logs first?
KubernetesSSH times out connecting to an EC2 instance. Ping also fails. But nc -zv <ip> 22 succeeds. What does this tell you?
Linux & NetworkingFour ways to access an EC2 instance when SSH is broken. List them in order of preference and explain the tradeoff.
Linux & NetworkingA client wants a 50/50 multi-cloud split "for cost savings." How would you push back or reframe this conversation?
Cloud Cost OptimizationWhat is an EC2 Launch Template and why does it matter for ASG migrations?
Cloud Cost OptimizationHow do you migrate 170 EC2 instances from Intel to AMD with zero unplanned downtime?
Cloud Cost OptimizationWhen is it appropriate to stop EC2 instances to save cost, and when is it not?
Cloud Cost OptimizationA production incident shows CPU, RAM, and disk all reporting as "healthy" in monitoring, yet a specific service is clearly struggling. What's a diagnostic angle that basic resource monitoring might miss?
Structured DebuggingWhat is a cluster autoscaler, and why does a Kubernetes cluster need one?
Cloud Cost OptimizationWhat is an S3 (or GCS) lifecycle policy, and what problem does it solve?
Cloud Cost OptimizationA client asks you to confirm that reducing a log bucket's retention period from 30 to 7 days won't affect their compliance posture. How would you verify and communicate this?
Cloud Cost OptimizationWhy does this system avoid synchronous cross-region database writes?
Career & Interview StrategyHow does the CAP theorem relate to the multi-region architecture decisions described in this guide?
Career & Interview StrategyA prospective client is based in the European Union and has strict data-residency requirements. You're evaluating CastAI (US-only SaaS) versus Karpenter (open-source, in-cluster) for their EKS autoscaling needs. Walk through your decision process.
Kubernetes AutoscalingHow does Karpenter differ from a standard EKS managed node group with Cluster Autoscaler?
Cloud Cost OptimizationWhy does EKS node group migration require ~2β3 minutes of downtime, while ASG migration has zero downtime?
Cloud Cost OptimizationDesign a node pool strategy for an EKS cluster that needs to support both a real-time customer-facing calculation service and a large nightly Spark analytics job, using the principles discussed in this guide.
Kubernetes AutoscalingWhat are the three EKS security scanning tools covered and what does each focus on?
Cloud Cost OptimizationHow does Karpenter know which subnets and security groups to use when provisioning a new node?
Kubernetes AutoscalingA client refuses to let you make application-level logging changes, but wants logging costs reduced. What are your levers, and what are the trade-offs of each?
Cloud Cost OptimizationA client insists on continuing to run both a self-hosted Prometheus/Grafana stack and a cloud-native monitoring tool in parallel, citing team preference. How would you approach cost optimization given this constraint, rather than pushing for consolidation?
Cloud Cost OptimizationWhy should multi-region deployments or patching operations always be sequential rather than parallel?
CI/CD & AutomationWhy is rollback done in parallel across all regions while forward deployment is sequential?
CI/CD & AutomationA service status check that normally responds instantly is now taking 8 seconds to respond. What would you conclude, and what would you do next?
Structured DebuggingA junior engineer on your team says a system issue "just fixed itself" during a live troubleshooting session, and they can no longer reproduce a bug you were actively diagnosing together. What questions would you ask before concluding the issue is actually resolved?
Structured DebuggingA client has a $98K/month AWS bill. Walk me through how you'd approach reducing it.
Cloud Cost OptimizationWhat tags should every AWS resource have, and why?
Cloud Cost Optimization