AWS CLI Triage Cheat Sheet

The AWS CLI commands you actually run during an incident, in the order you run them — from 'which account am I in' to a quota that stopped a scale-out.

Cloud command reference

The 60-second sweep

Six questions. Answer all of them before forming a theory.

# 1. Who am I, in which account, in which region?
aws sts get-caller-identity
aws configure get region

# 2. Is it AWS, or is it us?  (Business/Enterprise support)
aws health describe-events --filter eventStatusCodes=open \
  --query 'events[].[service,region,eventTypeCode]' --output table

# 3. What changed in the last hour?
aws cloudtrail lookup-events --max-results 50 \
  --start-time "$(date -u -v-1H +%Y-%m-%dT%H:%M:%SZ 2>/dev/null || date -u -d '1 hour ago' +%Y-%m-%dT%H:%M:%SZ)" \
  --query 'Events[].[EventTime,Username,EventName]' --output table

# 4. Is the path open?  (definitive answer, no packets sent)
aws ec2 create-network-insights-path --source i-0abc --destination i-0def \
  --protocol tcp --destination-port 443

# 5. Are we throttled?   CloudWatch ThrottledRequests on the service
# 6. Are we at a quota?
aws service-quotas list-service-quotas --service-code ec2 --output table

Step 1 resolves more reports than it has any right to: a stale AWS_PROFILE, a leftover AWS_ACCESS_KEY_ID in the shell, or the wrong region in one terminal.

Identity and permissions

# Will this call be allowed? Ask before you try.
aws iam simulate-principal-policy \
  --policy-source-arn arn:aws:iam::111122223333:role/app \
  --action-names s3:GetObject \
  --resource-arns arn:aws:s3:::bucket/key \
  --query 'EvaluationResults[].[EvalDecision,MatchedStatements[].SourcePolicyId]'

# What does this role actually hold?
aws iam list-attached-role-policies --role-name app
aws iam list-role-policies --role-name app

# Which SCPs apply to this account, and what is above it?
aws organizations list-policies-for-target --target-id 111122223333 \
  --filter SERVICE_CONTROL_POLICY
aws organizations list-parents --child-id 111122223333

# Long-lived keys nobody rotated
aws iam generate-credential-report >/dev/null
aws iam get-credential-report --query Content --output text | base64 -d | cut -d, -f1,9,11

Reading an AccessDenied: the reason clause names the layer. no identity-based policy allows → IAM. explicit deny in a service control policy → the organization. explicit deny in a resource-based policy → the bucket, queue or key policy.

Network

# Which route table is this subnet really using?
aws ec2 describe-route-tables \
  --filters Name=association.subnet-id,Values=subnet-0abc \
  --query 'RouteTables[].Routes[].[DestinationCidrBlock,GatewayId,NatGatewayId,State]' \
  --output table

# Subnet capacity — the answer to a surprising number of "it won't scale"
aws ec2 describe-subnets --filters Name=vpc-id,Values=vpc-0abc \
  --query 'Subnets[].[SubnetId,AvailabilityZone,CidrBlock,AvailableIpAddressCount]' \
  --output table

# Security groups open to the world
aws ec2 describe-security-groups --query \
 "SecurityGroups[?IpPermissions[?IpRanges[?CidrIp=='0.0.0.0/0']]].[GroupId,GroupName]" \
 --output table

# What is actually attached to this security group?
aws ec2 describe-network-interfaces --filters Name=group-id,Values=sg-0abc \
  --query 'NetworkInterfaces[].[NetworkInterfaceId,Description,PrivateIpAddress]' --output table

# NAT gateways — one per AZ, or a single point of failure?
aws ec2 describe-nat-gateways \
  --query 'NatGateways[].[NatGatewayId,SubnetId,State]' --output table
SymptomLook at
Instance unreachable from internetPublic IP present? IGW route? SG? NACL both directions?
Private instance cannot reach internetNAT route in this AZ’s route table
Connection hangs with no errorMissing VPC endpoint, or an endpoint SG with no inbound 443
Works from one subnet, not anotherDifferent route table association
Intermittent DNS failures under load1,024 packets/s per ENI to the .2 resolver — cache locally

EC2 and Auto Scaling

# Status checks: system = AWS's host, instance = your OS
aws ec2 describe-instance-status --instance-ids i-0abc \
  --query 'InstanceStatuses[].[SystemStatus.Status,InstanceStatus.Status]'

# The boot log, without any network
aws ec2 get-console-output --instance-id i-0abc --output text | tail -50

# In without SSH, without a bastion, without an inbound rule
aws ssm start-session --target i-0abc

# Why did the ASG not scale?
aws autoscaling describe-scaling-activities --auto-scaling-group-name app-asg \
  --max-items 10 --query 'Activities[].[StartTime,StatusCode,StatusMessage]' --output table

# Why is the target group unhealthy?
aws elbv2 describe-target-health --target-group-arn arn:aws:elasticloadbalancing:... \
  --query 'TargetHealthDescriptions[].[Target.Id,TargetHealth.State,TargetHealth.Reason]' \
  --output table

Elb.InitialHealthChecking means it is still starting. Target.FailedHealthChecks means the application answered wrongly. Target.Timeout means it did not answer.

Storage and data

# Volumes billing for nothing
aws ec2 describe-volumes --filters Name=status,Values=available \
  --query 'Volumes[].[VolumeId,Size,CreateTime]' --output table

# Still gp2 — cheaper and usually faster as gp3
aws ec2 describe-volumes --filters Name=volume-type,Values=gp2 \
  --query 'Volumes[].[VolumeId,Size]' --output table
aws ec2 modify-volume --volume-id vol-0abc --volume-type gp3

# RDS: is it Multi-AZ, encrypted, and how far behind are the replicas?
aws rds describe-db-instances --query \
 'DBInstances[].[DBInstanceIdentifier,MultiAZ,StorageEncrypted,DBInstanceStatus]' --output table

# Log groups with no retention — always a longer list than expected
aws logs describe-log-groups \
  --query 'logGroups[?!not_null(retentionInDays)].[logGroupName,storedBytes]' --output table

EKS

aws eks describe-cluster --name prod \
  --query 'cluster.{status:status,version:version,health:health}'

aws eks describe-nodegroup --cluster-name prod --nodegroup-name apps \
  --query 'nodegroup.{status:status,health:health,scaling:scalingConfig}'

# Pod IP budget — max pods, then subnet addresses
kubectl get node NODE -o jsonpath='{.status.allocatable.pods}{"\n"}'
aws ec2 describe-subnets --subnet-ids subnet-0a subnet-0b \
  --query 'Subnets[].[SubnetId,AvailableIpAddressCount]' --output table

# IRSA from inside the pod
kubectl exec -it deploy/api -n ns -- env | grep AWS_
kubectl exec -it deploy/api -n ns -- aws sts get-caller-identity

No AWS_ROLE_ARN in the pod means the webhook did not fire — the pod predates the ServiceAccount annotation, or the names do not match. Variables present but get-caller-identity failing means the role’s trust policy sub condition.

Quotas and cost

aws service-quotas get-service-quota --service-code ec2 --quota-code L-1216C47A
aws service-quotas request-service-quota-increase --service-code ec2 \
  --quota-code L-1216C47A --desired-value 512

# Yesterday's spend by service
aws ce get-cost-and-usage --time-period Start=2026-08-01,End=2026-08-18 \
  --granularity DAILY --metrics UnblendedCost \
  --group-by Type=DIMENSION,Key=SERVICE \
  --query 'ResultsByTime[-1].Groups[?Metrics.UnblendedCost.Amount>`10`].[Keys[0],Metrics.UnblendedCost.Amount]' \
  --output table

Output control worth memorising

--query 'Reservations[].Instances[].[InstanceId,State.Name]'   # JMESPath
--output table|text|json|yaml
--no-cli-pager                       # stop the pager in scripts
--dry-run                            # EC2 mutations: check permission only
--profile prod --region eu-west-1    # never rely on the ambient default

--dry-run is the cheapest habit in this file: it answers “am I allowed to do this, in this account” without doing it.