Regions, AZs and the Blast Radius Model

Why every AWS design decision reduces to one question — what dies with this resource — and how the global, regional and zonal scopes answer it.

beginner 18 min lesson hands-on task included

Most AWS material starts with a service tour. That is the wrong order. Services are a catalogue; what actually determines whether your system stays up is which scope each resource lives in, because scope is what decides what fails together.

Every resource in AWS is global, regional, or zonal. There is no fourth option, and the answer for any given resource is not negotiable.


Topic 1: The Three Scopes

EVERY RESOURCE LIVES IN EXACTLY ONE OF THESE SCOPES GLOBAL — one instance for the whole account IAM · Route 53 · CloudFront · WAF (CF scope) · Organizations REGION — an independent copy of the service S3 · DynamoDB · SQS · EKS control plane AZ a separate power, cooling, network EC2 · EBS · RDS node subnet (never spans AZs) AZ b separate power, cooling, network EC2 · EBS · RDS node subnet (never spans AZs) AZ c separate power, cooling, network EC2 · EBS · RDS node subnet (never spans AZs) THE ONLY QUESTION THAT MATTERS: WHAT DIES WITH IT? An EBS volume dies with its AZ. An S3 bucket survives an AZ. IAM survives a region. Design is choosing which of those you depend on.
The nesting is the whole model. A zonal resource cannot survive its AZ, a regional resource cannot survive its region, and a global resource has no region to lose — but also no regional isolation to protect it from a bad change.

Global — one instance for the entire account. IAM, Route 53, CloudFront, Organizations, and WAF when scoped to CloudFront. Creating an IAM role “in eu-west-1” is not a thing; the console shows you a region selector out of habit, and it is ignored.

Regional — an independent copy of the service in each region, with no automatic relationship between copies. S3, DynamoDB, SQS, Lambda, the EKS control plane. An S3 bucket in eu-west-1 and one in us-east-1 are two unrelated buckets that happen to be reachable through one API surface. Nothing crosses regions unless you configured it to.

Zonal — pinned to one Availability Zone for its whole life. EC2 instances, EBS volumes, subnets, individual RDS nodes, NAT gateways. You cannot move them; you replace them.

A region is a geographic area with at least three AZs. An AZ is one or more physically separate data centres with independent power, cooling and networking, connected to the other AZs in the region by low-latency private fibre — single-digit milliseconds, which is close enough for synchronous replication and far enough that a flood in one does not reach the others.

The trap in AZ names. eu-west-1a is not the same physical zone in your account and mine. AWS maps zone names to physical zones per account to stop everyone piling into “a”. If you need to talk about the same physical zone across accounts, use the AZ ID (euw1-az1), which is stable:

aws ec2 describe-availability-zones \
  --query 'AvailabilityZones[].[ZoneName,ZoneId,State]' --output table

Topic 2: Control Plane and Data Plane

This distinction explains most AWS incidents that look impossible.

The control plane is the API that creates, changes and describes resources: RunInstances, CreateBucket, PutScalingPolicy. The data plane is the thing your traffic actually touches: the instance serving requests, the load balancer forwarding packets, S3 returning objects.

AWS designs data planes to be static-stable — to keep working with the configuration they already have, even when the control plane is degraded. That produces behaviour worth internalising:

During a control-plane impairmentWhat happens
Running EC2 instancesKeep running and keep serving
Launching new instancesMay fail — this is control plane
An ALB routing to healthy targetsKeeps routing
Auto Scaling replacing a failed instanceMay stall — control plane
Existing IAM role sessionsKeep working
Assuming a role for a new sessionMay fail

The operational consequence: a recovery plan that depends on launching new capacity depends on the control plane, and the control plane is exactly what degrades in a large event. Pre-provisioned capacity that is already running is more reliable than capacity you intend to create during the incident. This is why “warm standby” beats “we’ll scale up if it happens” for anything with a tight RTO.

It is also why us-east-1 deserves special caution: several global services have control-plane dependencies there. Your workload in Frankfurt can be healthy while your ability to change it is not.


Topic 3: Service Resilience Tiers

AWS services are not equally resilient, and the differences are not advertised on the console front page. Three tiers:

GLOBALLY RESILIENT     IAM, Route 53, CloudFront
                       Survives losing a region entirely.

REGIONALLY RESILIENT   S3, DynamoDB, SQS, Lambda, ALB, EKS control plane
                       Replicated across AZs inside one region.
                       Survives losing an AZ. Does NOT survive losing the region.

ZONALLY RESILIENT      EC2 instance, EBS volume, single-AZ RDS, NAT gateway
                       Survives a host failure at best. Dies with its AZ.
                       (AWS calls these "zonal services" — the word
                       "resilient" is generous.)

Read that as a checklist against your own architecture. The interesting cases are the ones people misfile:

  • A NAT gateway is zonal. One NAT gateway serving all three AZs means an AZ failure removes outbound internet for the other two AZs’ private subnets. One per AZ, with each private route table pointing at its own — this is the single most common single point of failure in an otherwise multi-AZ VPC.
  • An ALB is regional, but its nodes are zonal. You enable it per subnet; enabling only two of three AZs means a third of your capacity is unreachable through it.
  • An EBS volume is zonal and cannot be attached across AZs. A snapshot is regional. That asymmetry is your migration path.
  • RDS Multi-AZ is two zonal nodes with a regional endpoint — the failover is what makes it survive an AZ, not the database being magically regional.

Topic 4: Choosing a Region

Four inputs, in this order:

  1. Data residency and law. GDPR, sector regulation, or a contract clause can eliminate every other consideration. This is a hard constraint, not a preference — and it decides the question before latency ever gets a vote.
  2. Latency to users. Roughly 1 ms per 100 km of fibre, plus real-world routing overhead. Frankfurt to Sydney is ~250 ms round trip and no amount of tuning fixes that; a CDN or a second region does.
  3. Service availability. New services launch in a handful of regions first. Check before you design, not after: aws ec2 describe-regions tells you regions; the service’s own docs tell you where it exists.
  4. Cost. Prices differ per region by a meaningful margin for the same instance. us-east-1 is usually cheapest and sa-east-1 and some APAC regions notably are not.

The pattern that keeps multi-region sane is one home region per workload that owns the writes, with other regions serving reads or standing by. Systems that let any region write to shared state acquire conflict resolution as a permanent tax, and conflict resolution is where correctness bugs live. Pick the home region deliberately and write it down.


Topic 5: Local Zones, Wavelength and Edge

Three extensions to the model, worth knowing exist so you can recognise them and mostly not use them:

  • Local Zones — a subset of services physically closer to a metro area, attached to a parent region. For single-digit-millisecond requirements only. Fewer services, higher prices.
  • Wavelength Zones — inside telco 5G networks, for mobile edge workloads.
  • Edge locations / regional edge caches — hundreds of PoPs used by CloudFront, Global Accelerator and S3 Transfer Acceleration. Not somewhere you deploy an instance; somewhere your content is cached or your TCP connection is terminated early.

The useful default: two or three AZs in one region, with a CDN for global reach. Reach for a second region when an actual requirement — an RTO you cannot otherwise hit, or a residency rule — demands it, because a second region roughly doubles your operational surface and every deployment, secret rotation and schema migration now happens twice.


Topic 6: Reading Scope Off a Resource

The ARN tells you the scope without any documentation:

arn:aws:iam::111122223333:role/deploy
                 ↑ no region  → GLOBAL

arn:aws:s3:::my-bucket
           ↑ no region, no account → globally unique NAME, regional resource

arn:aws:ec2:eu-west-1:111122223333:instance/i-0abc
            ↑ region, no AZ → the AZ is in the resource's own attributes

arn:aws:dynamodb:eu-west-1:111122223333:table/orders
                  ↑ region → REGIONAL

Format: arn:partition:service:region:account-id:resource. An empty region field means global. The partition matters more than people expect — aws-cn for China and aws-us-gov for GovCloud are separate partitions, and an ARN from one is meaningless in the other.

Try it yourself: run this and read the output as a scope inventory rather than a resource list.

# Zonal resources — every one of these dies with its AZ
aws ec2 describe-instances \
  --query 'Reservations[].Instances[].[InstanceId,Placement.AvailabilityZone,State.Name]' \
  --output table
aws ec2 describe-nat-gateways \
  --query 'NatGateways[].[NatGatewayId,SubnetId,State]' --output table
aws ec2 describe-volumes \
  --query 'Volumes[].[VolumeId,AvailabilityZone,Size]' --output table

If the NAT gateway list has one row and the instance list has three AZs, you have found a single point of failure in under a minute.

Common mistake: treating “multi-AZ” as a property you enable rather than a property you verify. Three subnets in three AZs with all the instances in one of them is a single-AZ deployment with extra billing complexity. Count what is actually running, per AZ, before you claim resilience — and remember that AWS will happily let an Auto Scaling group drift into one zone if the other subnets are out of IP addresses.