Every autoscaler is configured with three numbers, and most estates set them by feel. The gap between a felt number and a derived one is where a large share of Kubernetes overspend lives.
Topic 1: The Three Numbers
Minimum capacity — nodes running when the cluster starts. Sized to run the platform’s own components (DNS, networking agents, metrics) plus your most critical services. Below this the cluster does not function.
Desired capacity — steady state for normal operating load, excluding peak. Identify the critical processes and size for typical concurrent usage. This is the number that sets your baseline cost.
Maximum capacity — the ceiling. The most nuanced of the three, and the one most often wrong.
How not to set the maximum: “set it high, we can afford headroom.” That produces chronic overprovisioning and removes the ceiling’s only real function, which is to bound a runaway.
Topic 2: Deriving Maximum Capacity
A worked example from a banking platform, showing that the number comes from the business, not from infrastructure:
| Step | Question | Result |
|---|---|---|
| 1. Absolute upper bound | How many customers exist? | 50,000 |
| 2. Trend factor | What is the observed peak simultaneous session count from production logs? | 20,000 |
| 3. Growth factor | Customer base grew 30k → 50k (~1.67×). Scale the observed peak proportionally | ~33,000 |
| 4. Buffer | Add 20% for planning-to-implementation lag and business growth | ~40,000 |
| 5. Convert to infrastructure | Load test to find how many nodes serve 40,000 concurrent sessions | e.g. 20 nodes |
Three principles fall out of this:
- Never size for the absolute maximum. All 50,000 customers will not log in simultaneously. Sizing for it guarantees permanent, massive overprovisioning.
- Always add the 20% buffer. Capacity planning to implementation takes months, and the business grows in between.
- The inputs come from two teams. The application team supplies resource demand per session; business stakeholders supply upcoming campaigns, regulatory events and seasonal peaks. Infrastructure cannot produce this number alone.
For seasonal businesses — retail, anything with a tax or fiscal cycle, education — the seasonal peak must be an explicit input, and you need at least a year of history to see it. Sizing to a twelve-week average and meeting a seasonal peak are different exercises.
Topic 3: The Problem This Creates
You size for peak. Peak happens a handful of times a year. The rest of the time roughly 30% of provisioned capacity sits idle and billing.
That is not a configuration mistake — it is the structural consequence of static capacity with a peak-derived ceiling. It is also the precise problem the next lesson’s tooling exists to solve.
Topic 4: Where Classic Node Autoscaling Falls Short
The traditional approach scales a predefined node group: a fixed instance type, with min, desired and max configured by hand.
| Problem | Classic autoscaler | Why it costs money |
|---|---|---|
| Provisioning speed | 3–5 minutes to add a node via the scaling group | Slow response means you over-provision to compensate |
| Fixed instance type | Scales the type you pre-selected | A pod needing 2 vCPU may land on a node sized for 8 |
| Overprovisioning | ~30% idle at off-peak with peak-derived config | Direct, continuous waste |
| Management overhead | Engineers tune min/max/desired per node group | Tuning decays as workloads change |
The core limitation: it scales within a pool you defined in advance. It cannot ask “what is the cheapest machine that fits this pod?” because the answer was fixed when you created the node group.
A clarification worth making, since it is commonly misstated: the classic cluster autoscaler is not built into Kubernetes. It also requires installation. It is simply older, more established, and the default recommendation for longer.
Topic 5: Requests Are the Real Currency
Before any autoscaler can be efficient, the workloads must declare their needs honestly. The scheduler places pods on requests, not on actual usage — so a pod requesting 6 GB and using 1 GB consumes 6 GB of schedulable capacity.
resources:
requests:
cpu: "1"
memory: "2Gi"
limits:
cpu: "2"
memory: "4Gi"
Two consequences that matter for cost:
- Requests drive node size. Inflated requests force the autoscaler to provision bigger nodes than the workload needs. This is the single largest source of hidden cluster waste, and no node-level tool can detect it — a point the tooling lesson returns to.
- Requests equal to limits gives a guaranteed quality class. The pod will not be evicted under node pressure. That is the right choice for critical services and an expensive choice for everything else, because it reserves capacity that cannot be shared.
Workloads with no requests declared are worse still: the scheduler cannot reason about them at all, and bin-packing degrades to guesswork.
Topic 6: Autoscaling Is the Second Line of Defence
A framing worth adopting deliberately:
In a mature production environment you should already know your workload patterns by the time you reach production. Heavy versus lightweight, predictable versus spiky — that should have been identified in planning and validated in lower environments. An autoscaler exists to absorb genuinely unexpected load beyond what was planned for.
Treating the autoscaler as the first line of defence — provisioning nothing deliberately and letting the scaler work it out — produces a cluster whose cost profile nobody designed and nobody can predict.
Try it yourself: Look up the max node count configured on one of your node groups and ask where the number came from. If nobody can answer, derive it using the five-step method and compare. The two figures are rarely close.
Common mistake: Setting maximum capacity high “for safety”. The maximum is not a safety feature — it is a spending ceiling. Sized for the absolute theoretical peak, it stops bounding anything and simply permits unlimited scaling on a bad day.