Karpenter inverts the model from the previous lesson. Instead of scaling a pool you defined in advance, it reads the pending pod’s requests and provisions the cheapest instance that satisfies them — then removes it again when the work is done.
Topic 1: What It Actually Does
- Watches for pods stuck in
Pendingbecause no node can take them. - Reads the pod’s
resources.requests. - Selects the cheapest instance type from the set you permitted that satisfies the request.
- Calls the cloud API directly — bypassing scaling groups entirely.
- When load drops, bin-packs pods onto fewer nodes and terminates the remainder.
Provisioning lands in roughly 30–60 seconds against 3–5 minutes for scaling-group-based approaches, because there is no scaling group in the path.
What it does not replace: pod-level autoscaling. Horizontal and vertical pod autoscalers operate on pods; Karpenter operates on nodes. They are complementary layers, and you will usually run both.
Scope: per cluster. Ten clusters means ten installations. There is no cross-cluster coordination.
Topic 2: The Six-Step Setup
1 — Install via Helm, typically into the system namespace:
helm upgrade --install karpenter oci://public.ecr.aws/karpenter/karpenter \
--version "${KARPENTER_VERSION}" \
--namespace kube-system \
--set "settings.clusterName=${CLUSTER_NAME}" \
--set "settings.clusterEndpoint=${CLUSTER_ENDPOINT}" \
--set "settings.interruptionQueue=${INTERRUPTION_QUEUE}" \
--set serviceAccount.create=false
The interruption queue is what lets Karpenter react gracefully to reclamation notices on interruptible capacity — configure it if you intend to use Spot at all.
2 — Create the IAM role. Karpenter needs to create and terminate instances, describe them, pass roles, and read the interruption queue. Bind it to the service account through your cluster’s identity mechanism.
3 and 4 — Tag subnets and security groups with a discovery tag:
aws ec2 create-tags --resources <subnet-ids> \
--tags Key="karpenter.sh/discovery",Value="${CLUSTER_NAME}"
This is how Karpenter finds where it is allowed to launch. Without these tags it cannot provision anything, and the failure mode is silent pending pods — worth checking first when nothing happens.
5 — Define the node template:
apiVersion: karpenter.k8s.aws/v1beta1
kind: EC2NodeClass
metadata:
name: default
spec:
amiFamily: AL2023
role: "KarpenterNodeRole-${CLUSTER_NAME}"
subnetSelectorTerms:
- tags: { karpenter.sh/discovery: "${CLUSTER_NAME}" }
securityGroupSelectorTerms:
- tags: { karpenter.sh/discovery: "${CLUSTER_NAME}" }
6 — Define scheduling constraints:
apiVersion: karpenter.sh/v1beta1
kind: NodePool
metadata:
name: default
spec:
template:
spec:
nodeClassRef: { name: default }
requirements:
- key: "karpenter.sh/capacity-type"
operator: In
values: ["spot", "on-demand"]
- key: "kubernetes.io/arch"
operator: In
values: ["amd64", "arm64"]
- key: "karpenter.k8s.aws/instance-category"
operator: In
values: ["c", "m", "r"]
limits:
cpu: 1000
memory: 1000Gi
disruption:
consolidationPolicy: WhenUnderutilized
consolidateAfter: 30s
expireAfter: 720h
Do all of this through infrastructure as code. The CLI walkthrough is for understanding; the tagging in particular is exactly the kind of out-of-band state that breaks a cluster nobody can rebuild.
Topic 3: The Two Resources
EC2NodeClass is what nodes can look like — the image family, the node role, and which subnets and security groups are eligible. It is the equivalent of a launch template.
NodePool is how Karpenter schedules — the constraints, the ceiling, and the disruption policy.
Three fields on the NodePool carry most of the cost behaviour:
requirements— the wider the permitted set, the more freedom Karpenter has to find something cheap. Include a mix of capacity types and a mix of sizes. A pool restricted to one instance type gives the optimiser nothing to optimise.limits— a hard ceiling on total CPU and memory. This is your runaway protection, and it is the one field people forget.disruption— when to consolidate, and how long a node may live.
Topic 4: Consolidation
The primary off-peak saving mechanism, and the thing classic autoscaling could not do well.
When pods finish or scale down, Karpenter identifies underutilised nodes, cordons and drains them, reschedules the pods onto fewer nodes, and terminates the empties.
disruption:
consolidationPolicy: WhenUnderutilized
consolidateAfter: 30s
expireAfter: 720h # force replacement after 30 days
expireAfter is underrated: it guarantees nodes are periodically replaced, which is how you get image and kernel updates rolled in without a separate patching exercise.
Observed behaviour worth calibrating on: scale-up is fast — five pending pods took a three-node cluster to eight within one to two minutes, and the controller chose Spot capacity by itself based on the declared requirements. Scale-down is deliberately slower, taking around five minutes to settle, because consolidation waits to confirm the reduction is real rather than thrashing on a momentary dip.
The controller log is your troubleshooting entry point. During normal provisioning it emits info entries. Warning or error entries mean it failed to provision, which is why pods stay Pending — check there before checking anything else, and check the discovery tags first when it reports it cannot find a subnet.
Topic 5: Bin-Packing and Placement
Karpenter can only pack well if pods declare honestly. Requests are the input to every placement decision.
Isolate interruptible capacity with taints and tolerations so only workloads that tolerate reclamation land there:
tolerations:
- key: "karpenter.sh/capacity-type"
operator: "Equal"
value: "spot"
effect: "NoSchedule"
Core controllers and stateful services stay on stable capacity; ephemeral and batch workloads take the cheap capacity.
Topic 6: Node Pool Design
Should you pre-design pools for specific workload classes, or let one broad pool handle everything?
The trade-off is real. The more you constrain Karpenter with per-workload rules, the more you substitute your own upfront judgement for its dynamic optimisation. That helps for predictable workloads and reduces the benefit for genuinely unpredictable ones.
Patterns that hold up in production:
- Interruptible pool for tolerant workloads — batch, background processing — against an on-demand pool for business-critical, user-facing services. The most common and most defensible split.
- Dedicated GPU pools, kept separate from general CPU workloads.
Both are implemented with labels and taints on the pool, matched by selectors and tolerations on the workloads.
There is no universal formula for pool segmentation. It is a judgement call requiring input from application and business teams about which workloads tolerate what. Anyone offering a fixed answer has not asked enough questions about the workloads.
Handling image updates:
A genuine operational wrinkle. With static node groups you can blue-green: build the new image, stand up a new pool, drain the old. Karpenter provisions on demand, so there is no fixed set of nodes to roll.
Two workable strategies: run two NodePools with different node classes and shift workloads between them with labels, or use expireAfter so nodes are cycled onto the current image within a bounded window. Neither is as clean as a blue-green node group, and this is an honest rough edge rather than a solved problem.
Try it yourself: Apply a NodePool that permits only one instance type, deploy a mixed workload, and note what gets provisioned. Then widen it to eight types across two capacity types and repeat. The difference in what Karpenter selects is the whole argument for a permissive requirements block.
Common mistake: Omitting limits from the NodePool. Without a ceiling, a runaway workload — a crash-looping deployment with high requests, a misconfigured job — can provision nodes until you notice on the bill. The ceiling costs nothing when it is not being hit.