Foundations → Engagement

Cloud Cost Roadmap

Five stages, each gated by a capability rather than a lesson count. Cost work is the one discipline where doing it badly costs more than not doing it at all — an outage caused by a saving is a net loss — so every stage ends with a check you can either pass or cannot.

12
Lessons
5h
Total
5
Stages
5
Cost layers

Orderings that are not optional

Most lessons can be taken in any order within a stage. These four cannot — and the third one is the only mistake in this path you genuinely cannot undo.

  • Tagging before every other lesson Without attribution you cannot report a saving, target automation safely, or ask anyone to change behaviour.
  • A 90-day baseline before any right-sizing Sizing against a snapshot is guesswork. The baseline is also what tells you when not to resize at all.
  • Every other optimization before buying a commitment Each one reduces the spend a commitment is sized against. Buy first and you lock the pre-optimization baseline in for up to three years.
  • Resource requests before cluster cost tooling A cost platform will correctly report enormous waste that a free in-cluster right-sizer would have shown you for nothing.

Stage 1 — Foundations

~1h

Find out where the money actually goes before touching anything

3 lessons
Prerequisite

Console access to one cloud account and permission to read its billing data.

Everything in this stage is measurement. Nothing here saves a penny directly — and skipping it is why most cost work stalls after the first easy win.

You will be able to
  • ▸Read a bill as a diagnostic signal rather than a number to reduce
  • ▸Rank spend by layer and let the data pick the starting point, even when it sends you somewhere unfamiliar
  • ▸Design a tag schema and state what each tag unlocks downstream
  • ▸Build an inventory that doubles as the scope boundary for the whole engagement
  • ▸Collect a 90-day baseline and name the four situations where it cannot yet be trusted
  • ▸Enable the native toolchain on day one so recommenders have history when you need them
Mastery check

Given an unfamiliar account, you produce a layer-ranked cost analysis and a tag-coverage percentage inside a day — and you can name the layer you would attack first with a reason that is not "compute is usually biggest".

Stage 2 — Compute

~1h

The largest layer, and the one that can cause an outage

3 lessons
Prerequisite

Stage 1. A baseline of at least 90 days for anything you intend to resize.

You will be able to
  • ▸Apply the PRC framework so cost is never optimised at the expense of performance or reliability
  • ▸Size against P99 with a 20–33% buffer, and explain why sizing to the average causes incidents
  • ▸Migrate CPU architecture for 15–40% savings across all three instance categories
  • ▸Predict which migration is trivial and which is hard — the difficulty inverts between the two hops
  • ▸Schedule non-production shutdown with a tag filter that makes production structurally unreachable
  • ▸Place interruptible capacity in production safely, with the six safeguards that make it survivable
  • ▸Size a commitment against a baseline floor rather than total spend
Mastery check

You can hand someone a migration plan for a 170-instance fleet with a rollback path per category, and defend why the commitment purchase is scheduled last rather than first.

Stage 3 — Kubernetes Cost

~1h

Where declared demand and real demand diverge

3 lessons
Prerequisite

Stage 2, and working knowledge of pods, requests and the scheduler.

Cluster cost is mostly a requests problem wearing a node costume. Nearly everything in this stage traces back to pods asking for more than they use.

You will be able to
  • ▸Derive maximum capacity from business numbers instead of a round figure someone typed
  • ▸Explain why pool-based autoscaling structurally over-provisions by roughly a third
  • ▸Run demand-based provisioning with node pools permissive enough to actually optimise
  • ▸Set the ceiling that bounds a runaway, and read the controller log when nothing provisions
  • ▸Explain why an autoscaler cannot tell you a resource request was wrong
  • ▸Apply a five-criterion tool framework where data residency can disqualify a tool outright
Mastery check

Shown a cluster bill, you can separate waste caused by node provisioning from waste caused by inflated requests — and you know which tool addresses which, and which one you do not need to buy.

Stage 4 — Beyond Compute

~0.5h

The layers nobody checks, where waste accumulates unowned

2 lessons
Prerequisite

Stage 1. Independent of stages 2 and 3 — take it earlier if your billing data points here.

You will be able to
  • ▸Clear unattached volumes, snapshot sprawl and idle addresses — pure waste with no trade-off
  • ▸Set object lifecycle policies without stranding data you need during an incident
  • ▸Find the network charges that map to nothing you deliberately created
  • ▸Separate log volume from log retention and pull the safer lever first
  • ▸Distinguish compliance-fixed retention from negotiable retention before proposing anything
  • ▸Explain why filtering and restoration are not symmetric operations
Mastery check

You can produce a currency figure for pure-waste storage and networking in an hour, and you never promise a log "rollback" that cannot actually restore anything.

Stage 5 — Capstone Project

~1h

Run the whole thing, and make the savings defensible

1 lesson
Prerequisite

Stages 1–4.

You will be able to
  • ▸Sequence an engagement risk-ascending, so early wins fund the harder conversations
  • ▸Rate every plan item by both projected saving and implementation risk
  • ▸Validate in a lower environment before propagating, and know when pre-production is not one
  • ▸Report with per-resource, per-day evidence rather than a headline percentage
  • ▸Recognise when one change unlocks several others and reprioritise accordingly
  • ▸Name the engagement constraints — change window, dual-run tolerance, code-change mandate, ownership — in week one
Mastery check

You can hand a stakeholder four artifacts — inventory, layer analysis, risk-rated plan, implementation log — and every claimed saving survives the question "compared to what, and what else changed?"

One thing this path insists on

Cost is one pillar of three. Every sizing decision balances performance and reliability alongside it, and a resource that is genuinely under-provisioned should get bigger during a cost engagement — costing more, and being the correct call. A programme that causes an incident has not saved anything.