Most AWS material starts with a service tour. That is the wrong order. Services are a catalogue; what actually determines whether your system stays up is which scope each resource lives in, because scope is what decides what fails together.
Every resource in AWS is global, regional, or zonal. There is no fourth option, and the answer for any given resource is not negotiable.
Topic 1: The Three Scopes
Global — one instance for the entire account. IAM, Route 53, CloudFront, Organizations, and WAF when scoped to CloudFront. Creating an IAM role “in eu-west-1” is not a thing; the console shows you a region selector out of habit, and it is ignored.
Regional — an independent copy of the service in each region, with no automatic relationship between copies. S3, DynamoDB, SQS, Lambda, the EKS control plane. An S3 bucket in eu-west-1 and one in us-east-1 are two unrelated buckets that happen to be reachable through one API surface. Nothing crosses regions unless you configured it to.
Zonal — pinned to one Availability Zone for its whole life. EC2 instances, EBS volumes, subnets, individual RDS nodes, NAT gateways. You cannot move them; you replace them.
A region is a geographic area with at least three AZs. An AZ is one or more physically separate data centres with independent power, cooling and networking, connected to the other AZs in the region by low-latency private fibre — single-digit milliseconds, which is close enough for synchronous replication and far enough that a flood in one does not reach the others.
The trap in AZ names. eu-west-1a is not the same physical zone in your account and mine. AWS maps zone names to physical zones per account to stop everyone piling into “a”. If you need to talk about the same physical zone across accounts, use the AZ ID (euw1-az1), which is stable:
aws ec2 describe-availability-zones \
--query 'AvailabilityZones[].[ZoneName,ZoneId,State]' --output table
Topic 2: Control Plane and Data Plane
This distinction explains most AWS incidents that look impossible.
The control plane is the API that creates, changes and describes resources: RunInstances, CreateBucket, PutScalingPolicy. The data plane is the thing your traffic actually touches: the instance serving requests, the load balancer forwarding packets, S3 returning objects.
AWS designs data planes to be static-stable — to keep working with the configuration they already have, even when the control plane is degraded. That produces behaviour worth internalising:
| During a control-plane impairment | What happens |
|---|---|
| Running EC2 instances | Keep running and keep serving |
| Launching new instances | May fail — this is control plane |
| An ALB routing to healthy targets | Keeps routing |
| Auto Scaling replacing a failed instance | May stall — control plane |
| Existing IAM role sessions | Keep working |
| Assuming a role for a new session | May fail |
The operational consequence: a recovery plan that depends on launching new capacity depends on the control plane, and the control plane is exactly what degrades in a large event. Pre-provisioned capacity that is already running is more reliable than capacity you intend to create during the incident. This is why “warm standby” beats “we’ll scale up if it happens” for anything with a tight RTO.
It is also why us-east-1 deserves special caution: several global services have control-plane dependencies there. Your workload in Frankfurt can be healthy while your ability to change it is not.
Topic 3: Service Resilience Tiers
AWS services are not equally resilient, and the differences are not advertised on the console front page. Three tiers:
GLOBALLY RESILIENT IAM, Route 53, CloudFront
Survives losing a region entirely.
REGIONALLY RESILIENT S3, DynamoDB, SQS, Lambda, ALB, EKS control plane
Replicated across AZs inside one region.
Survives losing an AZ. Does NOT survive losing the region.
ZONALLY RESILIENT EC2 instance, EBS volume, single-AZ RDS, NAT gateway
Survives a host failure at best. Dies with its AZ.
(AWS calls these "zonal services" — the word
"resilient" is generous.)
Read that as a checklist against your own architecture. The interesting cases are the ones people misfile:
- A NAT gateway is zonal. One NAT gateway serving all three AZs means an AZ failure removes outbound internet for the other two AZs’ private subnets. One per AZ, with each private route table pointing at its own — this is the single most common single point of failure in an otherwise multi-AZ VPC.
- An ALB is regional, but its nodes are zonal. You enable it per subnet; enabling only two of three AZs means a third of your capacity is unreachable through it.
- An EBS volume is zonal and cannot be attached across AZs. A snapshot is regional. That asymmetry is your migration path.
- RDS Multi-AZ is two zonal nodes with a regional endpoint — the failover is what makes it survive an AZ, not the database being magically regional.
Topic 4: Choosing a Region
Four inputs, in this order:
- Data residency and law. GDPR, sector regulation, or a contract clause can eliminate every other consideration. This is a hard constraint, not a preference — and it decides the question before latency ever gets a vote.
- Latency to users. Roughly 1 ms per 100 km of fibre, plus real-world routing overhead. Frankfurt to Sydney is ~250 ms round trip and no amount of tuning fixes that; a CDN or a second region does.
- Service availability. New services launch in a handful of regions first. Check before you design, not after:
aws ec2 describe-regionstells you regions; the service’s own docs tell you where it exists. - Cost. Prices differ per region by a meaningful margin for the same instance.
us-east-1is usually cheapest andsa-east-1and some APAC regions notably are not.
The pattern that keeps multi-region sane is one home region per workload that owns the writes, with other regions serving reads or standing by. Systems that let any region write to shared state acquire conflict resolution as a permanent tax, and conflict resolution is where correctness bugs live. Pick the home region deliberately and write it down.
Topic 5: Local Zones, Wavelength and Edge
Three extensions to the model, worth knowing exist so you can recognise them and mostly not use them:
- Local Zones — a subset of services physically closer to a metro area, attached to a parent region. For single-digit-millisecond requirements only. Fewer services, higher prices.
- Wavelength Zones — inside telco 5G networks, for mobile edge workloads.
- Edge locations / regional edge caches — hundreds of PoPs used by CloudFront, Global Accelerator and S3 Transfer Acceleration. Not somewhere you deploy an instance; somewhere your content is cached or your TCP connection is terminated early.
The useful default: two or three AZs in one region, with a CDN for global reach. Reach for a second region when an actual requirement — an RTO you cannot otherwise hit, or a residency rule — demands it, because a second region roughly doubles your operational surface and every deployment, secret rotation and schema migration now happens twice.
Topic 6: Reading Scope Off a Resource
The ARN tells you the scope without any documentation:
arn:aws:iam::111122223333:role/deploy
↑ no region → GLOBAL
arn:aws:s3:::my-bucket
↑ no region, no account → globally unique NAME, regional resource
arn:aws:ec2:eu-west-1:111122223333:instance/i-0abc
↑ region, no AZ → the AZ is in the resource's own attributes
arn:aws:dynamodb:eu-west-1:111122223333:table/orders
↑ region → REGIONAL
Format: arn:partition:service:region:account-id:resource. An empty region field means global. The partition matters more than people expect — aws-cn for China and aws-us-gov for GovCloud are separate partitions, and an ARN from one is meaningless in the other.
Try it yourself: run this and read the output as a scope inventory rather than a resource list.
# Zonal resources — every one of these dies with its AZ
aws ec2 describe-instances \
--query 'Reservations[].Instances[].[InstanceId,Placement.AvailabilityZone,State.Name]' \
--output table
aws ec2 describe-nat-gateways \
--query 'NatGateways[].[NatGatewayId,SubnetId,State]' --output table
aws ec2 describe-volumes \
--query 'Volumes[].[VolumeId,AvailabilityZone,Size]' --output table
If the NAT gateway list has one row and the instance list has three AZs, you have found a single point of failure in under a minute.
Common mistake: treating “multi-AZ” as a property you enable rather than a property you verify. Three subnets in three AZs with all the instances in one of them is a single-AZ deployment with extra billing complexity. Count what is actually running, per AZ, before you claim resilience — and remember that AWS will happily let an Auto Scaling group drift into one zone if the other subnets are out of IP addresses.