Right-sizing is the fastest-payback optimization there is, and the one most likely to cause an incident. Both facts come from the same source: it changes the capacity a running workload has available to it.
Topic 1: PRC — Performance, Reliability, Cost
Every sizing decision balances three pillars. Cost is one of them, not the only one.
| Pillar | Means | Cost of ignoring it |
|---|---|---|
| P — Performance | The workload meets its latency and throughput objectives under expected load | Degradation, breached SLOs |
| R — Reliability | It absorbs traffic spikes and failure modes without falling over | Incidents from capacity starvation |
| C — Cost | Spend is minimised subject to the above | Budget overrun |
The rule: never optimise C at the expense of P or R. The goal is the smallest size that still satisfies performance and reliability with a safety margin — not the smallest size that runs.
This framing is worth stating explicitly to stakeholders at the start of an engagement, because it sets the expectation that some resources will not be resized, and one or two might get bigger.
Topic 2: Size Against P99, Not Average
The single most consequential technical point in this lesson.
A service averaging 20% CPU sounds like an obvious downsize. If its P99 is 85%, it is correctly sized and downsizing it will cause an incident at the next peak. The average describes the quiet hours; the P99 describes the moment that matters.
Collect per instance, over the baseline window:
- CPU — average, P95, P99
- Memory — average, P95, P99 (usually needs an agent installed; frequently missing, and its absence is itself a finding)
- Disk IOPS and throughput for storage-bound workloads
- Network throughput
Then size for the P99 plus headroom.
Topic 3: The Buffer Rule
Add 20–33% headroom above observed peak. Never size to exactly what the metrics showed.
The buffer absorbs three things the historical data cannot show you:
- Traffic growth between measurement and implementation — engagements take weeks, businesses keep growing.
- Unmeasured spikes shorter than your metric resolution. A five-minute sample interval hides a thirty-second saturation event completely.
- Degraded-mode load. When one instance in a pool fails, its traffic redistributes to the survivors. Size so that redistribution does not cascade.
That third point is the one most often forgotten, and it is the one that turns a single instance failure into an outage.
Topic 4: The Procedure
1 — Document the current state. Metrics, and the performance numbers you will compare against. Without a before, “it seems fine” is the only available verdict afterwards.
2 — Consult the recommender, but treat it as input rather than instruction. It cannot see the quarterly batch job or the launch next month.
3 — Change one variable. Size, or architecture, or storage type — not several at once. If something degrades you need to know which change caused it.
4 — Monitor for 72 hours minimum. Watch latency, error rate, memory pressure and out-of-memory events. Seventy-two hours covers a weekday cycle and at least one nightly batch window.
5 — If it degrades, go up one size — not straight back to the original. The original was probably too big. One step up from the new size is usually the right answer, and jumping back to the start discards what you just learned.
6 — Record the saving with evidence from the cost explorer, grouped by day.
Topic 5: When Not to Right-Size
Four situations where the correct action is to wait:
| Situation | Why |
|---|---|
| Uncertain or rapidly growing traffic | The headroom assumption expires before you finish |
| Migration in progress between providers | vCPU and memory performance ratios differ between platforms; resizing mid-move compounds two variables and you cannot attribute a regression to either |
| Within 30 days of a major feature release | The traffic pattern may have materially changed; the baseline describes a system that no longer exists |
| Within 30 days of provisioning | Not enough data to distinguish signal from noise |
Deferring is a legitimate, defensible outcome. “We are not right-sizing these twelve instances until after the launch, and here is why” is a better engagement note than a saving that gets reverted after an incident.
Topic 6: Right-Sizing Beyond Instances
The same discipline applies wherever capacity is declared:
- Managed databases — instance class, and storage autoscaling settings.
- Serverless functions — memory allocation drives both cost and CPU share. Over-allocating is expensive; under-allocating can be more expensive, because the function runs longer.
- Caches — node type and count.
- Kubernetes pod requests — the biggest source of hidden waste in a cluster, and the subject of stage 3.
That last one matters more than it looks. In a cluster, a pod requesting six times the memory it uses does not just waste that memory — it forces the autoscaler to provision nodes large enough to satisfy a request nobody needed.
Try it yourself: Pull average and P99 CPU for ten instances and sort by the gap between them. Large gaps mean spiky workloads that need headroom; small gaps mean steady workloads that can be sized tightly. That one column changes how you treat each instance.
Common mistake: Right-sizing everything in a single change window because the recommendations were generated together. Twenty simultaneous resizes means twenty candidate causes when latency moves. Batch them by blast radius and leave time to observe between batches.