AWS CLI Triage Cheat Sheet
The AWS CLI commands you actually run during an incident, in the order you run them — from 'which account am I in' to a quota that stopped a scale-out.
The 60-second sweep
Six questions. Answer all of them before forming a theory.
# 1. Who am I, in which account, in which region?
aws sts get-caller-identity
aws configure get region
# 2. Is it AWS, or is it us? (Business/Enterprise support)
aws health describe-events --filter eventStatusCodes=open \
--query 'events[].[service,region,eventTypeCode]' --output table
# 3. What changed in the last hour?
aws cloudtrail lookup-events --max-results 50 \
--start-time "$(date -u -v-1H +%Y-%m-%dT%H:%M:%SZ 2>/dev/null || date -u -d '1 hour ago' +%Y-%m-%dT%H:%M:%SZ)" \
--query 'Events[].[EventTime,Username,EventName]' --output table
# 4. Is the path open? (definitive answer, no packets sent)
aws ec2 create-network-insights-path --source i-0abc --destination i-0def \
--protocol tcp --destination-port 443
# 5. Are we throttled? CloudWatch ThrottledRequests on the service
# 6. Are we at a quota?
aws service-quotas list-service-quotas --service-code ec2 --output table
Step 1 resolves more reports than it has any right to: a stale AWS_PROFILE,
a leftover AWS_ACCESS_KEY_ID in the shell, or the wrong region in one terminal.
Identity and permissions
# Will this call be allowed? Ask before you try.
aws iam simulate-principal-policy \
--policy-source-arn arn:aws:iam::111122223333:role/app \
--action-names s3:GetObject \
--resource-arns arn:aws:s3:::bucket/key \
--query 'EvaluationResults[].[EvalDecision,MatchedStatements[].SourcePolicyId]'
# What does this role actually hold?
aws iam list-attached-role-policies --role-name app
aws iam list-role-policies --role-name app
# Which SCPs apply to this account, and what is above it?
aws organizations list-policies-for-target --target-id 111122223333 \
--filter SERVICE_CONTROL_POLICY
aws organizations list-parents --child-id 111122223333
# Long-lived keys nobody rotated
aws iam generate-credential-report >/dev/null
aws iam get-credential-report --query Content --output text | base64 -d | cut -d, -f1,9,11
Reading an AccessDenied: the reason clause names the layer. no identity-based policy allows → IAM. explicit deny in a service control policy → the organization. explicit deny in a resource-based policy → the bucket, queue or key policy.
Network
# Which route table is this subnet really using?
aws ec2 describe-route-tables \
--filters Name=association.subnet-id,Values=subnet-0abc \
--query 'RouteTables[].Routes[].[DestinationCidrBlock,GatewayId,NatGatewayId,State]' \
--output table
# Subnet capacity — the answer to a surprising number of "it won't scale"
aws ec2 describe-subnets --filters Name=vpc-id,Values=vpc-0abc \
--query 'Subnets[].[SubnetId,AvailabilityZone,CidrBlock,AvailableIpAddressCount]' \
--output table
# Security groups open to the world
aws ec2 describe-security-groups --query \
"SecurityGroups[?IpPermissions[?IpRanges[?CidrIp=='0.0.0.0/0']]].[GroupId,GroupName]" \
--output table
# What is actually attached to this security group?
aws ec2 describe-network-interfaces --filters Name=group-id,Values=sg-0abc \
--query 'NetworkInterfaces[].[NetworkInterfaceId,Description,PrivateIpAddress]' --output table
# NAT gateways — one per AZ, or a single point of failure?
aws ec2 describe-nat-gateways \
--query 'NatGateways[].[NatGatewayId,SubnetId,State]' --output table
| Symptom | Look at |
|---|---|
| Instance unreachable from internet | Public IP present? IGW route? SG? NACL both directions? |
| Private instance cannot reach internet | NAT route in this AZ’s route table |
| Connection hangs with no error | Missing VPC endpoint, or an endpoint SG with no inbound 443 |
| Works from one subnet, not another | Different route table association |
| Intermittent DNS failures under load | 1,024 packets/s per ENI to the .2 resolver — cache locally |
EC2 and Auto Scaling
# Status checks: system = AWS's host, instance = your OS
aws ec2 describe-instance-status --instance-ids i-0abc \
--query 'InstanceStatuses[].[SystemStatus.Status,InstanceStatus.Status]'
# The boot log, without any network
aws ec2 get-console-output --instance-id i-0abc --output text | tail -50
# In without SSH, without a bastion, without an inbound rule
aws ssm start-session --target i-0abc
# Why did the ASG not scale?
aws autoscaling describe-scaling-activities --auto-scaling-group-name app-asg \
--max-items 10 --query 'Activities[].[StartTime,StatusCode,StatusMessage]' --output table
# Why is the target group unhealthy?
aws elbv2 describe-target-health --target-group-arn arn:aws:elasticloadbalancing:... \
--query 'TargetHealthDescriptions[].[Target.Id,TargetHealth.State,TargetHealth.Reason]' \
--output table
Elb.InitialHealthChecking means it is still starting. Target.FailedHealthChecks
means the application answered wrongly. Target.Timeout means it did not answer.
Storage and data
# Volumes billing for nothing
aws ec2 describe-volumes --filters Name=status,Values=available \
--query 'Volumes[].[VolumeId,Size,CreateTime]' --output table
# Still gp2 — cheaper and usually faster as gp3
aws ec2 describe-volumes --filters Name=volume-type,Values=gp2 \
--query 'Volumes[].[VolumeId,Size]' --output table
aws ec2 modify-volume --volume-id vol-0abc --volume-type gp3
# RDS: is it Multi-AZ, encrypted, and how far behind are the replicas?
aws rds describe-db-instances --query \
'DBInstances[].[DBInstanceIdentifier,MultiAZ,StorageEncrypted,DBInstanceStatus]' --output table
# Log groups with no retention — always a longer list than expected
aws logs describe-log-groups \
--query 'logGroups[?!not_null(retentionInDays)].[logGroupName,storedBytes]' --output table
EKS
aws eks describe-cluster --name prod \
--query 'cluster.{status:status,version:version,health:health}'
aws eks describe-nodegroup --cluster-name prod --nodegroup-name apps \
--query 'nodegroup.{status:status,health:health,scaling:scalingConfig}'
# Pod IP budget — max pods, then subnet addresses
kubectl get node NODE -o jsonpath='{.status.allocatable.pods}{"\n"}'
aws ec2 describe-subnets --subnet-ids subnet-0a subnet-0b \
--query 'Subnets[].[SubnetId,AvailableIpAddressCount]' --output table
# IRSA from inside the pod
kubectl exec -it deploy/api -n ns -- env | grep AWS_
kubectl exec -it deploy/api -n ns -- aws sts get-caller-identity
No AWS_ROLE_ARN in the pod means the webhook did not fire — the pod predates
the ServiceAccount annotation, or the names do not match. Variables present but
get-caller-identity failing means the role’s trust policy sub condition.
Quotas and cost
aws service-quotas get-service-quota --service-code ec2 --quota-code L-1216C47A
aws service-quotas request-service-quota-increase --service-code ec2 \
--quota-code L-1216C47A --desired-value 512
# Yesterday's spend by service
aws ce get-cost-and-usage --time-period Start=2026-08-01,End=2026-08-18 \
--granularity DAILY --metrics UnblendedCost \
--group-by Type=DIMENSION,Key=SERVICE \
--query 'ResultsByTime[-1].Groups[?Metrics.UnblendedCost.Amount>`10`].[Keys[0],Metrics.UnblendedCost.Amount]' \
--output table
Output control worth memorising
--query 'Reservations[].Instances[].[InstanceId,State.Name]' # JMESPath
--output table|text|json|yaml
--no-cli-pager # stop the pager in scripts
--dry-run # EC2 mutations: check permission only
--profile prod --region eu-west-1 # never rely on the ambient default
--dry-run is the cheapest habit in this file: it answers “am I allowed to do
this, in this account” without doing it.