Capacity Planning & the Limits of Cluster Autoscaler

Deriving min, desired and max capacity from business numbers rather than guesswork, the 20% buffer rule, and the three structural problems that make classic node autoscaling overprovision by design.

advanced 20 min lesson hands-on task included

Every autoscaler is configured with three numbers, and most estates set them by feel. The gap between a felt number and a derived one is where a large share of Kubernetes overspend lives.


Topic 1: The Three Numbers

Minimum capacity — nodes running when the cluster starts. Sized to run the platform’s own components (DNS, networking agents, metrics) plus your most critical services. Below this the cluster does not function.

Desired capacity — steady state for normal operating load, excluding peak. Identify the critical processes and size for typical concurrent usage. This is the number that sets your baseline cost.

Maximum capacity — the ceiling. The most nuanced of the three, and the one most often wrong.

How not to set the maximum: “set it high, we can afford headroom.” That produces chronic overprovisioning and removes the ceiling’s only real function, which is to bound a runaway.


Topic 2: Deriving Maximum Capacity

A worked example from a banking platform, showing that the number comes from the business, not from infrastructure:

StepQuestionResult
1. Absolute upper boundHow many customers exist?50,000
2. Trend factorWhat is the observed peak simultaneous session count from production logs?20,000
3. Growth factorCustomer base grew 30k → 50k (~1.67×). Scale the observed peak proportionally~33,000
4. BufferAdd 20% for planning-to-implementation lag and business growth~40,000
5. Convert to infrastructureLoad test to find how many nodes serve 40,000 concurrent sessionse.g. 20 nodes

Three principles fall out of this:

  • Never size for the absolute maximum. All 50,000 customers will not log in simultaneously. Sizing for it guarantees permanent, massive overprovisioning.
  • Always add the 20% buffer. Capacity planning to implementation takes months, and the business grows in between.
  • The inputs come from two teams. The application team supplies resource demand per session; business stakeholders supply upcoming campaigns, regulatory events and seasonal peaks. Infrastructure cannot produce this number alone.

For seasonal businesses — retail, anything with a tax or fiscal cycle, education — the seasonal peak must be an explicit input, and you need at least a year of history to see it. Sizing to a twelve-week average and meeting a seasonal peak are different exercises.


Topic 3: The Problem This Creates

You size for peak. Peak happens a handful of times a year. The rest of the time roughly 30% of provisioned capacity sits idle and billing.

That is not a configuration mistake — it is the structural consequence of static capacity with a peak-derived ceiling. It is also the precise problem the next lesson’s tooling exists to solve.


Topic 4: Where Classic Node Autoscaling Falls Short

The traditional approach scales a predefined node group: a fixed instance type, with min, desired and max configured by hand.

ProblemClassic autoscalerWhy it costs money
Provisioning speed3–5 minutes to add a node via the scaling groupSlow response means you over-provision to compensate
Fixed instance typeScales the type you pre-selectedA pod needing 2 vCPU may land on a node sized for 8
Overprovisioning~30% idle at off-peak with peak-derived configDirect, continuous waste
Management overheadEngineers tune min/max/desired per node groupTuning decays as workloads change
Pod goes Pending — what happens next pool-based detect pending scaling group adds a node — 3–5 minutes fixed type — may be far too big demand-based read pod requests 30–60s cheapest instance that fits — right-sized by construction Slow provisioning is why pool-based clusters over-provision: you hold spare capacity to hide the wait. Consolidation then reclaims idle nodes when load drops — the off-peak saving pools cannot make.
The provisioning gap is why pool-based clusters over-provision structurally: you hold idle capacity to hide a three-to-five-minute wait. Remove the wait and the idle capacity stops being necessary.

The core limitation: it scales within a pool you defined in advance. It cannot ask “what is the cheapest machine that fits this pod?” because the answer was fixed when you created the node group.

A clarification worth making, since it is commonly misstated: the classic cluster autoscaler is not built into Kubernetes. It also requires installation. It is simply older, more established, and the default recommendation for longer.


Topic 5: Requests Are the Real Currency

Before any autoscaler can be efficient, the workloads must declare their needs honestly. The scheduler places pods on requests, not on actual usage — so a pod requesting 6 GB and using 1 GB consumes 6 GB of schedulable capacity.

resources:
  requests:
    cpu: "1"
    memory: "2Gi"
  limits:
    cpu: "2"
    memory: "4Gi"

Two consequences that matter for cost:

  • Requests drive node size. Inflated requests force the autoscaler to provision bigger nodes than the workload needs. This is the single largest source of hidden cluster waste, and no node-level tool can detect it — a point the tooling lesson returns to.
  • Requests equal to limits gives a guaranteed quality class. The pod will not be evicted under node pressure. That is the right choice for critical services and an expensive choice for everything else, because it reserves capacity that cannot be shared.

Workloads with no requests declared are worse still: the scheduler cannot reason about them at all, and bin-packing degrades to guesswork.


Topic 6: Autoscaling Is the Second Line of Defence

A framing worth adopting deliberately:

In a mature production environment you should already know your workload patterns by the time you reach production. Heavy versus lightweight, predictable versus spiky — that should have been identified in planning and validated in lower environments. An autoscaler exists to absorb genuinely unexpected load beyond what was planned for.

Treating the autoscaler as the first line of defence — provisioning nothing deliberately and letting the scaler work it out — produces a cluster whose cost profile nobody designed and nobody can predict.


Try it yourself: Look up the max node count configured on one of your node groups and ask where the number came from. If nobody can answer, derive it using the five-step method and compare. The two figures are rarely close.

Common mistake: Setting maximum capacity high “for safety”. The maximum is not a safety feature — it is a spending ceiling. Sized for the absolute theoretical peak, it stops bounding anything and simply permits unlimited scaling on a bad day.