Cloud Cost Roadmap
Five stages, each gated by a capability rather than a lesson count. Cost work is the one discipline where doing it badly costs more than not doing it at all — an outage caused by a saving is a net loss — so every stage ends with a check you can either pass or cannot.
Orderings that are not optional
Most lessons can be taken in any order within a stage. These four cannot — and the third one is the only mistake in this path you genuinely cannot undo.
- Tagging before every other lesson Without attribution you cannot report a saving, target automation safely, or ask anyone to change behaviour.
- A 90-day baseline before any right-sizing Sizing against a snapshot is guesswork. The baseline is also what tells you when not to resize at all.
- Every other optimization before buying a commitment Each one reduces the spend a commitment is sized against. Buy first and you lock the pre-optimization baseline in for up to three years.
- Resource requests before cluster cost tooling A cost platform will correctly report enormous waste that a free in-cluster right-sizer would have shown you for nothing.
Stage 1 — Foundations
~1hFind out where the money actually goes before touching anything
Console access to one cloud account and permission to read its billing data.
Everything in this stage is measurement. Nothing here saves a penny directly — and skipping it is why most cost work stalls after the first easy win.
- ▸Read a bill as a diagnostic signal rather than a number to reduce
- ▸Rank spend by layer and let the data pick the starting point, even when it sends you somewhere unfamiliar
- ▸Design a tag schema and state what each tag unlocks downstream
- ▸Build an inventory that doubles as the scope boundary for the whole engagement
- ▸Collect a 90-day baseline and name the four situations where it cannot yet be trusted
- ▸Enable the native toolchain on day one so recommenders have history when you need them
Given an unfamiliar account, you produce a layer-ranked cost analysis and a tag-coverage percentage inside a day — and you can name the layer you would attack first with a reason that is not "compute is usually biggest".
Stage 2 — Compute
~1hThe largest layer, and the one that can cause an outage
Stage 1. A baseline of at least 90 days for anything you intend to resize.
- ▸Apply the PRC framework so cost is never optimised at the expense of performance or reliability
- ▸Size against P99 with a 20–33% buffer, and explain why sizing to the average causes incidents
- ▸Migrate CPU architecture for 15–40% savings across all three instance categories
- ▸Predict which migration is trivial and which is hard — the difficulty inverts between the two hops
- ▸Schedule non-production shutdown with a tag filter that makes production structurally unreachable
- ▸Place interruptible capacity in production safely, with the six safeguards that make it survivable
- ▸Size a commitment against a baseline floor rather than total spend
You can hand someone a migration plan for a 170-instance fleet with a rollback path per category, and defend why the commitment purchase is scheduled last rather than first.
Stage 3 — Kubernetes Cost
~1hWhere declared demand and real demand diverge
Stage 2, and working knowledge of pods, requests and the scheduler.
Cluster cost is mostly a requests problem wearing a node costume. Nearly everything in this stage traces back to pods asking for more than they use.
- ▸Derive maximum capacity from business numbers instead of a round figure someone typed
- ▸Explain why pool-based autoscaling structurally over-provisions by roughly a third
- ▸Run demand-based provisioning with node pools permissive enough to actually optimise
- ▸Set the ceiling that bounds a runaway, and read the controller log when nothing provisions
- ▸Explain why an autoscaler cannot tell you a resource request was wrong
- ▸Apply a five-criterion tool framework where data residency can disqualify a tool outright
Shown a cluster bill, you can separate waste caused by node provisioning from waste caused by inflated requests — and you know which tool addresses which, and which one you do not need to buy.
Stage 4 — Beyond Compute
~0.5hThe layers nobody checks, where waste accumulates unowned
Stage 1. Independent of stages 2 and 3 — take it earlier if your billing data points here.
- ▸Clear unattached volumes, snapshot sprawl and idle addresses — pure waste with no trade-off
- ▸Set object lifecycle policies without stranding data you need during an incident
- ▸Find the network charges that map to nothing you deliberately created
- ▸Separate log volume from log retention and pull the safer lever first
- ▸Distinguish compliance-fixed retention from negotiable retention before proposing anything
- ▸Explain why filtering and restoration are not symmetric operations
You can produce a currency figure for pure-waste storage and networking in an hour, and you never promise a log "rollback" that cannot actually restore anything.
Stage 5 — Capstone Project
~1hRun the whole thing, and make the savings defensible
Stages 1–4.
- ▸Sequence an engagement risk-ascending, so early wins fund the harder conversations
- ▸Rate every plan item by both projected saving and implementation risk
- ▸Validate in a lower environment before propagating, and know when pre-production is not one
- ▸Report with per-resource, per-day evidence rather than a headline percentage
- ▸Recognise when one change unlocks several others and reprioritise accordingly
- ▸Name the engagement constraints — change window, dual-run tolerance, code-change mandate, ownership — in week one
You can hand a stakeholder four artifacts — inventory, layer analysis, risk-rated plan, implementation log — and every claimed saving survives the question "compared to what, and what else changed?"
One thing this path insists on
Cost is one pillar of three. Every sizing decision balances performance and reliability alongside it, and a resource that is genuinely under-provisioned should get bigger during a cost engagement — costing more, and being the correct call. A programme that causes an incident has not saved anything.