The AWS Failure Playbook

A fixed six-step sweep for any AWS incident, the four error classes and what each actually means, and the quota and throttling failures that present as random.

advanced 22 min lesson hands-on task included

Under pressure, people debug by intuition, and intuition goes to the thing they changed last or the thing they understand best. A fixed sequence beats intuition because it eliminates whole categories of cause in a known order, and because it is repeatable when you are tired.


Topic 1: The Sweep

FIRST SIX QUESTIONS — ANSWER ALL OF THEM BEFORE FORMING A THEORY 1 Who am I, in which account? aws sts get-caller-identity --query Arn the one people skip 2 Is it AWS or is it us? AWS Health Dashboard aws health describe-events 3 What changed in the last hour? cloudtrail lookup-events --max-results 50 4 Is the path even open? ec2 describe-security-groups Reachability Analyzer 5 Are we being throttled? CloudWatch ThrottledRequests 4xx vs 5xx on the API 6 Are we at a quota? service-quotas list-service-quotas trusted-advisor limit checks WHY STEP 1 IS FIRST A surprising share of "the resource is gone" reports are a stale profile pointed at the wrong account, or the wrong region in one shell.
Answer all six before forming a theory. Step 1 looks trivial and resolves a genuinely surprising share of reports — a stale profile or the wrong region in one shell.
# 1. Who am I, in which account, in which region?
aws sts get-caller-identity
echo "region: ${AWS_REGION:-$(aws configure get region)}"

# 2. Is it AWS, or is it us?
aws health describe-events --filter regions=eu-west-1,eventStatusCodes=open \
  --query 'events[].[service,eventTypeCode,startTime]' --output table
#    (Business/Enterprise support required; otherwise the Health Dashboard)

# 3. What changed in the last hour?
aws cloudtrail lookup-events --max-results 50 \
  --start-time "$(date -u -d '1 hour ago' +%Y-%m-%dT%H:%M:%SZ)" \
  --query 'Events[].[EventTime,Username,EventName]' --output table

# 4. Is the network path even open?
aws ec2 describe-security-groups --group-ids sg-0abc \
  --query 'SecurityGroups[].IpPermissions'
#    then Reachability Analyzer for the definitive answer

# 5. Are we being throttled?
#    CloudWatch: ThrottledRequests, 4xx vs 5xx on the service's own metrics

# 6. Are we at a quota?
aws service-quotas list-service-quotas --service-code ec2 \
  --query 'Quotas[?contains(QuotaName,`On-Demand`)].[QuotaName,Value]' --output table

Steps 2 and 3 are the highest-yield pair. Most incidents are either AWS having a problem or someone having changed something, and both are answerable in under a minute. Only when both come back clean is it worth digging into your own application.

Record what normal looks like before you need it. A sweep whose output you have never seen healthy is much harder to read under pressure.


Topic 2: The Four Error Classes

Almost every AWS error is one of four things, and each has a different first move.

AccessDenied — a permission problem, and the message says which kind. Read the reason clause: no identity-based policy allows points at IAM, with an explicit deny in a service control policy points at the organization, with an explicit deny in a resource-based policy points at the bucket or key policy. Do not add permissions until you know which. (Full evaluation order is in lesson 2.)

ThrottlingException / RequestLimitExceeded / 429 — you are calling an API faster than your allowance. The fix is client-side: exponential backoff with jitter, which the AWS SDKs implement by default and which custom code and shell loops usually do not. A while true; do aws ...; done loop is a throttling incident waiting for a busy day.

InsufficientInstanceCapacity — AWS does not have that instance type in that AZ right now. Not a quota; genuinely no capacity. Try another AZ, another instance type, or a different size. This is why diversified instance types in an ASG or Karpenter node pool are a resilience feature, not only a cost one.

LimitExceeded / QuotaExceeded — you hit a service quota. Different from capacity, and fixable in advance.

Two more that matter because they are so often misread:

  • 5xx from an AWS API means AWS’s side failed; retry with backoff. 4xx means your request was wrong; retrying it verbatim will fail forever, which is worth knowing before you build a retry loop.
  • A timeout with no error at all, on a call that normally works, is very often a missing VPC endpoint or a security group on an endpoint — the request goes nowhere and nothing answers.

Topic 3: Quotas

Nearly every AWS quantity has a limit, most are per region per account, and some are adjustable while others are hard.

The ones that bite first, in roughly the order people meet them:

VPCs per region                 5     (adjustable)
Elastic IPs per region          5     (adjustable — and the classic surprise)
Rules per security group       60     (adjustable)
Security groups per ENI         5     (hard)
Subnets per VPC               200     (adjustable)
Routes per route table         50     (adjustable to 1000)
On-Demand vCPUs per family      varies (adjustable — the one that stops a scale-out)
Lambda concurrent executions 1000     (adjustable)
# What is my actual limit, and what am I using?
aws service-quotas get-service-quota --service-code ec2 \
  --quota-code L-1216C47A   # Running On-Demand Standard instances

aws service-quotas request-service-quota-increase --service-code ec2 \
  --quota-code L-1216C47A --desired-value 512

Request increases before you need them. An increase can take hours or days, and it always arrives during the event you needed it for. Trusted Advisor reports usage against limits, and CloudWatch can alarm on quota utilisation directly — an alarm at 80% of a binding quota is one of the cheapest incidents you will ever prevent.

Quotas are per account, which is one more argument for the account-per-workload split: a load test that exhausts a quota in its own account does not stop production from scaling.


Topic 4: The Failures That Present as Random

Some failure modes look like flakiness and are entirely deterministic once you know what to measure.

DNS packet limit. 1,024 packets per second per ENI to the VPC resolver. A service resolving on every request with no caching hits it under load and gets intermittent resolution failures — which looks like a broken DNS server and is actually a rate limit. Cache locally; on Kubernetes, run NodeLocal DNS.

NAT gateway port exhaustion. 55,000 simultaneous connections per unique destination. Many instances talking to one endpoint through one NAT gateway can exhaust it. The signature is the ErrorPortAllocation metric being non-zero, and the symptom is connection failures that correlate with load and nothing else.

Burst credit exhaustion. t-family CPU credits, gp2 volume burst, EFS burst throughput. All present the same way: performance is fine, then abruptly is not, while utilisation metrics look unremarkable. The tell is a *BalanceCredit metric trending to zero over hours before the failure.

Connection reuse against an idle timeout. The ALB’s 60-second idle timeout versus an application keep-alive shorter than it produces intermittent 502s — the load balancer reuses a connection the application has just closed. Application keep-alive must exceed the load balancer’s idle timeout.

Cross-AZ everything. Not a failure, but the same shape of surprise: traffic that silently crosses AZs bills per gigabyte in both directions and adds latency. It appears as an unexplained data transfer line and an unexplained latency floor.

Eventual consistency in control planes. A role created and immediately assumed, a security group created and immediately referenced, an IAM policy attached and immediately relied on — all can fail for a few seconds. Automation that creates and uses in the same breath needs a retry, and this is why Terraform occasionally fails on the first apply and succeeds on the second.


Topic 5: Instance-Level Triage When SSH Is Not the Answer

When an instance is unreachable, the ordered options — most of which do not require the network to be working:

  1. EC2 status checks. System check failed means AWS’s host; stop-start moves you. Instance check failed means your OS; keep reading.
  2. Instance console output — the boot log, without any network. This is where a kernel panic, a failing fsck, or a full root filesystem announces itself:
aws ec2 get-console-output --instance-id i-0abc --output text | tail -50
aws ec2 get-console-screenshot --instance-id i-0abc --output text  # for a hung boot
  1. SSM Session Manager — works without SSH, without a bastion, and without an inbound rule, as long as the agent is running and can reach the SSM endpoints.
  2. SSM Run Command — execute a diagnostic script on many instances at once without connecting to any of them.
  3. EC2 Instance Connect Endpoint — SSH over an AWS-managed tunnel to a private instance.
  4. Detach the root volume, attach it to a working instance, and read or repair the filesystem. The last resort, and the one that recovers data when nothing else does.

The general lesson is that “I cannot SSH in” removes one option out of six, and the console output is usually faster than fixing SSH.


Topic 6: Writing It Down Afterwards

An incident that produces no artifact will happen again with the same duration.

During: one person coordinates, one scribe records timestamps, and the timeline is written as it happens rather than reconstructed. Reconstruction always loses the detail that mattered.

After: a blameless review answering four questions — what happened, what the impact was, why it took as long as it did to detect, and why it took as long as it did to fix. Separating detection from repair is what turns “we should be more careful” into “we need an alarm on this metric”.

The output is actions with owners. A review that produces a document and no changes is a document. Typical genuinely useful actions: an alarm that would have caught it twenty minutes earlier, a runbook entry for the fix, a quota increase, a guardrail that makes the trigger impossible, or a drill that would have surfaced the gap.

Try it yourself: run all six sweep steps against a healthy account and save the output. It takes two minutes now and turns the first two minutes of your next incident into pattern-matching instead of exploration.

Common mistake: starting an investigation in CloudTrail on a busy account. Tens of millions of events, no time bound, and forty minutes gone before the first useful fact. Narrow the window with a metric first — the alarm timestamp is your filter — and only then ask who did what. The order of the sweep is not arbitrary; each step makes the next one cheap.