EKS Cluster Anatomy and Node Capacity

What AWS runs and what you still own, the pod-per-node IP arithmetic that decides your subnet sizing, and the endpoint-access setting that can lock everyone out of a healthy cluster.

advanced 24 min lesson hands-on task included

EKS is Kubernetes with the control plane operated by AWS and everything else operated by you. The boundary is precise, and knowing exactly where it falls is what separates “the cluster is broken” from “our node group is broken”, which are different pages and different fixes.

The Kubernetes concepts themselves live in the Kubernetes path; this lesson is only about what AWS adds and takes away.


Topic 1: The Split

AWS-MANAGED VPC — you cannot see or ssh into any of this apiserver multi-AZ, autoscaled etcd backed up by AWS scheduler you never patch it controller-mgr $0.10/hour, flat cross-account ENIs in YOUR subnets YOUR VPC — everything you are actually responsible for managed node group ASG + AMI + drain on update Karpenter nodes no ASG, picks the shape Fargate profiles one pod per microVM add-ons VPC CNI · CoreDNS · kube-proxy kubelet → apiserver over the cluster endpoint · node IAM role · your AMI patching · your CNI IP budget THE IP MATH THAT BITES max pods ≈ (ENIs × (IPs per ENI − 1)) + 2 m5.large = 3 × 9 + 2 = 29 pods, no matter how idle the CPU is. ENDPOINT ACCESS IS A ONE-LINE OUTAGE Private-only endpoint + no VPN/bastion route = a cluster nobody can reach, including your CI.
AWS runs the top box and bills a flat hourly rate for it. Everything in the bottom box — nodes, AMIs, IP budget, add-on versions — is yours, and it is where every EKS incident actually lives.

AWS runs: the API server, etcd, the scheduler and controller manager, across at least two AZs, autoscaled, backed up, and patched. You cannot SSH to it, you cannot see the instances, and you cannot break it with a bad manifest. Flat rate, currently $0.10 per hour per cluster regardless of size.

You run: nodes, the AMIs on them, the add-ons, the IAM wiring, the networking, and every workload. Which means every EKS problem is one of: a node problem, an IP problem, an IAM problem, an add-on version problem, or an ordinary Kubernetes problem.

The control plane reaches your workloads through cross-account ENIs placed in the subnets you nominated at cluster creation. That is why the cluster’s subnets need enough free addresses, and why the security group on those ENIs matters — the API server talks to webhooks and to the kubelet through them.

The three node options:

Managed node groupKarpenterFargate
Backed byAn ASG AWS managesDirect EC2 API calls, no ASGOne microVM per pod
Chooses instance typeYou, from a listIt does, from your constraintsn/a
Scale-up latencyASG launch + join, minutesTypically under a minute~60s per pod
Node updatesManaged, with drainDrift-based replacementn/a
Best forSteady baseline capacityBursty, diverse, cost-sensitiveIsolation-sensitive, small, spiky

Most production clusters run a small managed node group for system add-ons — the things that must exist before anything can schedule — plus Karpenter for workload capacity. Karpenter’s advantage is that it picks the instance shape to fit the pending pods rather than scaling a fixed shape, which both packs better and scales faster. The trade is that it needs its own IAM, its own controller to keep alive, and node consolidation settings that you must understand before turning them on.


Topic 2: The IP Arithmetic

The AWS VPC CNI gives every pod a routable VPC address. This is why pod-to-pod traffic needs no overlay and security groups can apply to pods — and why address planning is a first-class EKS concern.

max pods per node ≈ (number of ENIs × (IPs per ENI − 1)) + 2

m5.large    3 ENIs × 10 IPs → (3 × 9) + 2 = 29
m5.xlarge   4 ENIs × 15 IPs → (4 × 14) + 2 = 58
m5.4xlarge  8 ENIs × 30 IPs → (8 × 29) + 2 = 234
t3.medium   3 ENIs ×  6 IPs → (3 ×  5) + 2 = 17

Verify rather than trusting the arithmetic:

kubectl get node <node> -o jsonpath='{.status.allocatable.pods}{"\n"}'

Two ceilings, two different Pending messages. Hitting the per-node limit gives you a pod that cannot schedule because no node has a free pod slot; hitting the subnet limit gives you failed to assign an IP address to container. The second is the one that gets misdiagnosed, because the message names the container runtime and the cause is a full subnet.

Three levers when you run out:

  • Prefix delegation — ENABLE_PREFIX_DELEGATION=true assigns /28 prefixes to ENIs rather than single addresses, raising an m5.large from 29 pods to 110. It consumes subnet space in blocks of 16, so it trades subnet efficiency for node density, and it requires Nitro instances.
  • Custom networking — pods get addresses from a secondary VPC CIDR (conventionally 100.64.0.0/16) via ENIConfig, leaving your primary space for nodes and everything else. This is the standard rescue for a cluster already in production.
  • Bigger subnets — the fix you can only apply before the cluster exists, which is why lesson 5 laboured the point.

The CNI also keeps a warm pool of addresses per node (WARM_IP_TARGET, MINIMUM_IP_TARGET). Defaults reserve a whole ENI’s worth, which on a large cluster is a meaningful fraction of your subnet held in reserve. Tuning it down reduces address pressure at the cost of slower pod starts during a burst.


Topic 3: Endpoint Access — the Configuration That Locks You Out

The cluster’s Kubernetes API endpoint has three modes:

PUBLIC              reachable from the internet, IAM-authenticated.
                    Restrict with publicAccessCidrs.
PUBLIC + PRIVATE    in-VPC traffic resolves privately, outside traffic
                    goes to the public endpoint. The usual choice.
PRIVATE ONLY        no public endpoint at all. Requires VPN, Direct Connect
                    or a bastion inside the VPC — for humans AND for CI.

Switching to private-only without a working in-VPC path is the classic self-inflicted EKS outage: workloads keep running, and nobody — including your pipeline — can change anything. Before flipping it, confirm that your CI runners are in the VPC or reachable through it, and that you have a break-glass path that does not depend on the change you are making.

Nodes also need to reach the endpoint. A fully private cluster needs interface endpoints for the services the kubelet and add-ons use: ec2, ecr.api, ecr.dkr, s3 (gateway), logs, sts, and elasticloadbalancing if you run the load balancer controller. Missing sts is a memorable one — nodes join fine and every IRSA-using pod fails to get credentials.


Topic 4: Add-ons and Versions

EKS add-ons are AWS-managed versions of the components a cluster needs. The four to know:

  • VPC CNI — pod networking, as above.
  • CoreDNS — cluster DNS. Two replicas by default, which is not enough for a busy cluster; scale it and consider NodeLocal DNS to keep the .2 resolver’s per-ENI packet limit out of your critical path.
  • kube-proxy — service routing. Its version must track the control plane’s.
  • EBS CSI driver — persistent volumes. It is not installed by default, and since Kubernetes 1.23 the in-tree provisioner is gone, so a cluster without it has PVCs that stay Pending forever with no obvious cause.
aws eks describe-addon-versions --addon-name vpc-cni \
  --kubernetes-version 1.31 --query 'addons[].addonVersions[].addonVersion'
aws eks update-addon --cluster-name prod --addon-name vpc-cni \
  --addon-version v1.19.0-eksbuild.1 --resolve-conflicts PRESERVE

Upgrades. EKS supports a rolling window of Kubernetes versions and moves clusters to extended support (at a higher price) when a version leaves standard support. The order is: control plane first, then add-ons, then nodes — and never skip a minor version. Before any upgrade, check deprecated API usage in your manifests; the API removals in a minor version are the thing that actually breaks workloads, not the upgrade mechanism.

Managed node group updates drain nodes respecting PodDisruptionBudgets, which means a PDB with minAvailable equal to the replica count blocks the drain indefinitely. That interaction is covered in depth in the Kubernetes path’s upgrade lesson, and it is the most common reason an EKS node group update appears to hang.


Topic 5: Load Balancers, Storage and the AWS Integrations

The AWS Load Balancer Controller turns Kubernetes objects into AWS resources: an Ingress becomes an ALB, a Service of type LoadBalancer with the right annotations becomes an NLB.

apiVersion: networking.k8s.io/v1
kind: Ingress
metadata:
  annotations:
    alb.ingress.kubernetes.io/scheme: internet-facing
    alb.ingress.kubernetes.io/target-type: ip          # direct to pod IPs
    alb.ingress.kubernetes.io/healthcheck-path: /healthz
    alb.ingress.kubernetes.io/certificate-arn: arn:aws:acm:...

target-type: ip sends traffic straight to pod IPs, skipping the kube-proxy hop that instance mode requires. It is faster, it preserves the client source IP more cleanly, and it consumes pod IPs from the same budget as everything else — which is the recurring theme of this lesson.

Storage: the EBS CSI driver for ReadWriteOnce volumes (zonal — a pod using one is pinned to that AZ, which interacts with scheduling in ways worth knowing before a StatefulSet is Pending), and the EFS CSI driver for ReadWriteMany across AZs.

Fargate deserves one honest paragraph: it removes node management entirely, isolates each pod in its own microVM, and cannot run DaemonSets, privileged pods, or anything needing host access. Logging requires the built-in log router rather than a node agent. It is excellent for a small number of isolation-sensitive workloads and a poor fit as a general-purpose default, mostly because per-pod pricing exceeds well-packed EC2 for steady load.


Topic 6: Cluster Triage, AWS-Side

When something is wrong, the AWS-side checks that are not in kubectl:

# Is the control plane actually healthy, and what version?
aws eks describe-cluster --name prod \
  --query 'cluster.{status:status,version:version,endpoint:endpoint,health:health}'

# Did the node group fail to scale, and why?
aws eks describe-nodegroup --cluster-name prod --nodegroup-name apps \
  --query 'nodegroup.{status:status,health:health,scaling:scalingConfig}'

# Are the subnets out of addresses? This is the answer more often than it should be.
aws ec2 describe-subnets --subnet-ids subnet-0a subnet-0b subnet-0c \
  --query 'Subnets[].[SubnetId,AvailableIpAddressCount]' --output table

# Control plane logs — off by default, and worth enabling before you need them
aws eks update-cluster-config --name prod \
  --logging '{"clusterLogging":[{"types":["api","audit","authenticator"],"enabled":true}]}'

Control plane logging being off by default is worth acting on today rather than during an incident: without the authenticator log there is no record of failed authentication attempts, and without audit there is no record of who deleted the deployment.

Try it yourself: fill a node to its pod ceiling with a scaled-up deployment of a tiny image. Watch the next pod stay Pending while the node reports plenty of free CPU and memory. Then read kubectl describe node and find the pods allocatable number that actually stopped you.

Common mistake: sizing an EKS cluster’s subnets from the expected node count. A three-node /24 looks generous until each node holds 58 pods, the CNI reserves a warm pool per node, a rolling deployment doubles pod count briefly, and the ALB controller adds targets by IP. The failure arrives during a deployment, at the worst possible time, and the only quick fix — a secondary CIDR with custom networking — requires recycling every node.