A VPC’s CIDR block cannot be changed after creation. You can add secondary CIDRs later, but you cannot shrink, renumber or move the primary — and neither can you change a subnet’s CIDR or its AZ. This is the one AWS decision where getting the arithmetic right on day one saves a migration on day four hundred.
Topic 1: Just Enough CIDR
A CIDR block is an address plus a prefix length: 10.20.0.0/16. The prefix says how many leading bits are fixed; the rest are yours.
/16 → 65,536 addresses 10.20.0.0 – 10.20.255.255
/20 → 4,096 addresses 10.20.0.0 – 10.20.15.255
/24 → 256 addresses 10.20.0.0 – 10.20.0.255
/28 → 16 addresses 10.20.0.0 – 10.20.0.15
Each +1 to the prefix halves the block. /24 → /25 = two halves.
Three notations worth having memorised because they appear in security groups and route tables constantly:
0.0.0.0/0— every address. In a route table it is the default route; in a security group it is “the entire internet”.10.20.0.0/32— exactly one address. How you allow a single host.10.20.0.0/16— a whole private range, the usual shape of a VPC.
Use RFC 1918 space — 10.0.0.0/8, 172.16.0.0/12, 192.168.0.0/16 — and pick from it deliberately. Avoid 172.17.0.0/16: that is Docker’s default bridge network, and a VPC that overlaps it produces containers that cannot reach half your infrastructure for reasons that take a day to find.
The rule that governs everything else: never overlap. Two VPCs with overlapping CIDRs cannot be peered, cannot both attach usefully to a Transit Gateway, and cannot both connect to the same on-premises network. Allocate ranges centrally, write them down, and treat the document as production infrastructure. A spreadsheet is fine; no spreadsheet is not.
Topic 2: Sizing a VPC
AWS allows /16 to /28 for a VPC. Practical guidance:
/16is the default answer. 65,536 addresses costs nothing, and running out is genuinely painful.- Reserve two
/16s per region per environment if your organisation is large enough to grow into it. Address space is free; renumbering is not. - Leave the second half of the VPC empty for expansion. Carving
10.20.0.0/17into subnets and leaving10.20.128.0/17untouched gives you room for a tier you have not thought of yet.
A layout that reviews well:
VPC 10.20.0.0/16
public-a 10.20.0.0/20 (4096) ALB nodes, NAT gateway
public-b 10.20.16.0/20 (4096)
public-c 10.20.32.0/20 (4096)
10.20.48.0/20 ← reserved for a 4th AZ
private-app-a 10.20.64.0/20 (4096) EC2, EKS nodes and pods
private-app-b 10.20.80.0/20 (4096)
private-app-c 10.20.96.0/20 (4096)
10.20.112.0/20 ← reserved
private-data-a 10.20.128.0/22 (1024) RDS, ElastiCache
private-data-b 10.20.132.0/22 (1024)
private-data-c 10.20.136.0/22 (1024)
10.20.144.0/20 … ← ~45% of the VPC unallocated
Public subnets can be small — they hold load balancer nodes and NAT gateways, not fleets. Data subnets can be small. The app tier is where you spend your addresses, and the next topic explains why it needs far more than the instance count suggests.
Topic 3: The Five Addresses You Do Not Get
In every subnet, AWS reserves five addresses:
| Address | Purpose |
|---|---|
.0 | Network address |
.1 | VPC router |
.2 | Amazon-provided DNS (the “VPC resolver”) |
.3 | Reserved for future use |
| last | Broadcast address (AWS does not support broadcast, but reserves it anyway) |
So a /24 gives 251 usable addresses, not 256. A /28 gives 11, and /28 is the smallest subnet AWS permits — which makes it a poor choice for anything except a tiny endpoint or a firewall subnet.
The .2 address is more important than it looks: it is the resolver every instance uses by default, and it also enforces a limit of 1,024 packets per second per network interface for DNS. A busy service with an aggressive resolver — no DNS caching, a short TTL, a client library that resolves per request — hits that ceiling and produces intermittent resolution failures that look exactly like a broken DNS server. The fix is caching on the instance (or NodeLocal DNS in Kubernetes), not a bigger subnet.
Topic 4: Why Kubernetes Changes the Arithmetic
The AWS VPC CNI gives every pod a real VPC IP address. Not an overlay address — a routable one from your subnet. This is excellent for observability and security groups, and brutal for capacity planning.
An m5.large has 3 ENIs × 10 IPv4 addresses each.
max pods ≈ (ENIs × (IPs per ENI − 1)) + 2 = (3 × 9) + 2 = 29
A /24 app subnet = 251 usable addresses
→ ~8 nodes' worth of pods, and that is before the CNI's
warm-IP pool takes a share up front.
Two independent ceilings, and either one will stop you:
- Per-node: the ENI/IP limit above. An
m5.largecaps at 29 pods no matter how idle its CPU is. This is why “the node has 80% free memory but pods are Pending” happens. - Per-subnet: total addresses across all nodes and pods in that subnet.
Mitigations, in the order you should reach for them:
- Size the app subnets for pods —
/20per AZ rather than/24. This is the cheap fix and it must happen before the cluster exists. - Add a secondary CIDR to the VPC (
100.64.0.0/16from the carrier-grade NAT space is the conventional choice) and run pods there withENIConfigcustom networking. This is the standard rescue for a cluster that has already run out. - Prefix delegation —
ENABLE_PREFIX_DELEGATION=trueassigns /28 prefixes to ENIs instead of individual IPs, raising per-node pod density dramatically on Nitro instances. It consumes subnet space in blocks, so it trades subnet efficiency for node density. - IPv6 — removes the address-scarcity problem entirely, and introduces a dual-stack complexity budget you must be ready to spend.
This is the most common self-inflicted EKS constraint, and it always shows up months later when someone scales a deployment, so it is worth over-provisioning subnet space at the start.
Topic 5: Secondary CIDRs, IPAM and IPv6
Secondary CIDR blocks can be added to an existing VPC (up to five, and they must not overlap with anything reachable). This is how you rescue a VPC that was sized too small. New subnets can use the new range; existing ones cannot move into it.
Amazon VPC IPAM is the service for organisations where “who owns 10.42.0.0/16” has become a real question. It allocates from pools, enforces non-overlap, and reports utilisation. Below roughly ten VPCs a documented spreadsheet is honestly fine; above it, IPAM stops being bureaucracy and starts preventing outages.
IPv6 in a VPC has one property worth internalising: there are no private IPv6 addresses in AWS — every IPv6 address is globally unique and potentially routable. Outbound-only access therefore needs an egress-only internet gateway, which is the IPv6 counterpart to a NAT gateway and, pleasingly, is free. AWS assigns you a fixed /56 per VPC and you carve /64s per subnet; the arithmetic is trivially generous, which is the point.
Topic 6: DNS Inside the VPC
Two VPC attributes control name resolution, and one of them defaults to off:
| Attribute | Default (custom VPC) | Effect |
|---|---|---|
enableDnsSupport | true | The .2 resolver answers queries at all |
enableDnsHostnames | false | Instances get public DNS names for their public IPs |
”My instance has no public DNS name” and “my private hosted zone does not resolve” are both usually enableDnsHostnames being false. Both attributes must be true for a private hosted zone to work at all:
aws ec2 describe-vpc-attribute --vpc-id vpc-0abc --attribute enableDnsHostnames
aws ec2 modify-vpc-attribute --vpc-id vpc-0abc --enable-dns-hostnames
DHCP option sets let you override the resolver and the search domain. Almost always the wrong move: replacing the Amazon resolver means losing resolution for VPC endpoints, private hosted zones and EC2 internal names unless you forward carefully. The supported way to reach on-premises DNS is Route 53 Resolver endpoints — outbound rules for your domains, inbound for their queries about yours — which keeps the .2 resolver in the path.
Try it yourself: create a /28 subnet and read its AvailableIpAddressCount. It says 11. Then launch an instance and read it again. That number, and how fast it falls in a Kubernetes subnet, is the whole capacity story.
Common mistake: sizing subnets from the current instance count. An ASG that scales, a rolling deployment that briefly doubles pod count, and a CNI warm pool all consume addresses simultaneously. A subnet at 85% utilisation on an ordinary Tuesday has no room for a deployment, and the failure surfaces as pods stuck in Pending with failed to assign an IP address to container — a message that names the network but gets blamed on Kubernetes.