Project: A Layered Cost Optimization Engagement

Run a complete engagement end to end — inventory and baseline, layer-by-layer analysis, a sequenced plan ordered by risk, implementation with verification, and the reporting that makes savings defensible.

advanced 45 min lesson hands-on task included

Everything in this module, assembled once. A cost engagement is not a list of optimizations — it is a sequence in which each step is individually verifiable and individually reversible, and where the reporting is good enough that somebody else can audit it a year later.


Topic 1: The Phases

Phase 1 — FOUNDATION        (week 1)
  ├─ Tag audit and remediation
  ├─ Full resource inventory  → audit trail
  ├─ Enable native cost tooling  → starts collecting
  └─ Begin 90-day baseline (runs in background)

Phase 2 — ANALYSIS          (week 1–2)
  ├─ Rank spend by layer
  ├─ Categorise: over / correctly / under-provisioned
  └─ Build the plan: saving + risk per item

Phase 3 — IMPLEMENTATION    (week 2+)
  ├─ Zero-risk waste first
  ├─ Low-risk changes second
  └─ Architectural changes last, with sign-off

Phase 4 — VERIFY & REPORT   (continuous)
  ├─ Per-change cost evidence, grouped by day
  ├─ Performance regression check
  └─ Weekly reporting cadence
Four phases — implementation runs risk-ascending 1 · FOUNDATION tags · inventory enable tooling start 90-day baseline 2 · ANALYSIS rank layers by spend categorise resources saving + risk per item 3 · IMPLEMENT validate in non-prod then propagate one variable at a time 4 · VERIFY per-day cost evidence performance check weekly reporting Order inside phase 3: zero-risk waste low risk architectural commitments — always last Every optimization above reduces the spend a commitment is sized against. Buy early and you lock the pre-optimization baseline in for up to three years. Anything that changes a failure domain needs written business sign-off before implementation.
The phases run left to right; inside phase 3 the work runs risk-ascending. Commitments sit at the far right of both orderings for the same reason — everything to their left changes the number they would be sized against.

Phase 1 finishing before phase 3 starts is what separates an engagement from a series of guesses.

Validate in a lower environment first, then propagate. Pre-production is the testbed: apply the change, observe the cost impact and the performance impact, and only then roll it to production. On engagements with tight change windows this is not optional — and note that some organisations treat pre-production with production-equivalent care, in which case you need a third environment to experiment in.


Topic 2: Phase 1 — Foundation

Tag audit. Report coverage as a percentage of resources with a complete tag set. Under 80% means every subsequent cost report is partly fiction. Remediate before analysing.

Inventory. One row per resource: type, size, state, tags, utilisation, cost. One sheet per category, consolidated into a single workbook.

The inventory is also your scope boundary. If the estate grows later and someone observes that costs did not fall, the snapshot defines what was in scope. Anything provisioned afterwards is demonstrably outside your work. This has saved more engagements than any optimization in it.

Enable the tooling on day one. Recommenders need history. Enabling them in week three wastes two weeks of collection.

Start the baseline. 90 days minimum, longer for seasonal workloads. It runs in the background while you do everything else.


Topic 3: Phase 2 — Analysis and the Plan

Rank the layers by actual spend. Do not assume compute leads — one estate had 44% in observability.

Then build a plan where every line has both a projected saving and a risk rating:

ItemLayerProjected savingRiskPrerequisite
Delete unattached volumesStorage$232/moNoneConfirm data unwanted
Release idle addressesNetwork$45/moNone—
Set log retention 30d → 14dObservability~$2,500/moLowTeam confirms need
Non-prod schedulingCompute~$900/moLowTag coverage complete
Volume type migrationStorage$180/moLow—
x86 vendor migration, 120 instancesCompute~$935/moLowImage backups
Private endpoints for object storageNetworkvariesMediumNetwork review
Right-size from baselineComputevariesMedium90-day baseline complete
ARM migrationCompute~$1,500/moMediumDependency scan clean
Cluster autoscaler → demand-based provisioningKubernetesvariesMediumRequests reviewed
Commitment purchaseCompute20–60% of baselineHighAll other work complete
Availability-zone consolidationNetworkvariesHighBusiness sign-off

Two ordering rules are non-negotiable:

Commitments come last. Every optimization above reduces the on-demand spend a commitment is sized against. Buy early and you lock in the pre-optimization baseline for up to three years — the one mistake in this module you genuinely cannot undo.

Anything that changes failure domains needs explicit business sign-off, in writing, before implementation.


Topic 4: Phase 3 — Implementation

Work risk-ascending. Early wins fund credibility for the harder conversations.

Every change gets the same treatment:

  1. Document the before state — metrics and cost.
  2. Confirm the rollback path and test it at least once.
  3. Change one variable.
  4. Monitor 48–72 hours.
  5. Verify cost in the explorer, grouped by day.
  6. Record the result.

Batch by blast radius. Twenty simultaneous resizes means twenty candidate causes when latency moves.

Clean up after yourself. Image backups taken for migration rollback cost money. Delete them once stable — usually one to two weeks. A cost engagement that leaves behind a pile of snapshots has partially undone itself, and it is a genuinely embarrassing finding in a review.


Topic 5: Phase 4 — Verification and Reporting

Per change: before and after cost grouped by day, plus a performance check confirming no regression. Both, every time. A saving with no performance evidence is an unverified claim.

Weekly: cumulative saving, items completed, items in flight, blockers.

A monthly bill total is not evidence of anything. Bills move for reasons unrelated to your work — traffic growth, a new service, a pricing change. Per-resource, per-day evidence attributes the change to the action.

Report what you did not do, and why. “We are not right-sizing these twelve instances until after the launch” and “these twelve instances were underprovisioned and we increased them” both belong in the report. They demonstrate the PRC framework was applied rather than cost being maximised in isolation.


Topic 6: The Constraints That Shape Everything

Four engagement constraints that decide what is even possible, and which you should establish in week one rather than discover in week four:

1 — The change window. Some businesses run continuously across regions and have no natural quiet period. A realistic constraint looks like 30–60 minutes, Saturdays only. That single fact rules out several optimizations and forces others to be staged.

2 — Willingness to dual-run. A client with a hard uptime requirement will often happily pay for old and new infrastructure simultaneously during a cutover, because stability outranks cost during the migration itself. Ask — it makes otherwise impossible migrations safe, and it is frequently offered rather than requested.

3 — Infrastructure-only versus code changes. Establish the mandate. Config, instance types, storage classes and retention are usually in scope by default. Anything needing application changes — beyond minor compatibility work like ARM-compatible libraries — needs a separate conversation and sign-off.

4 — Who owns what. The biggest risk to any cost engagement is the absence of a clear infrastructure owner. Where a platform grew organically under application developers, there is often no escalation path and no subject-matter owner to ask before touching a component. Any change made without understanding a component’s role can cause a business-impacting outage — and if there is nobody to ask, that risk is on you.

Tooling procurement is its own constraint. Introducing any third party — even a provider’s own advisory tooling behind a paid support tier — usually needs business approval. Budget time for it rather than assuming you can enable something on Tuesday.


Topic 7: An Alternative Layer Order

The default order in lesson 01 works down from compute. There is a defensible variant that runs network → compute → data, and it is worth knowing why.

The argument: fixing orphaned and wasteful network resources first avoids spending effort right-sizing compute that an architectural or network change is about to replace anyway. If you already suspect the topology is wrong, sequencing network first stops you optimising something twice.

Both orders are legitimate. What matters is that the order is chosen from the billing data and the known architectural work ahead, not inherited from a template.


Topic 8: Savings Cascade Through Layers

One change frequently unlocks several others, and the compounding is easy to miss when you evaluate items in isolation.

A worked example: a database carrying roughly 4 TB of a decade’s inactive user data. Cleaning it up cascaded:

  • Smaller volume → lower block storage cost
  • Smaller dataset → lower cross-region transfer cost for the daily backup copy
  • Smaller footprint → able to move off a premium database engine onto a standard one
  • Smaller engine → a further instance downsize on top

Four savings across three layers from one piece of data hygiene. When you rank the plan by projected saving, evaluate whether any item unlocks others — the unlocking item deserves priority even when its own direct saving looks modest.


Topic 9: Multi-Phase Engagements

Cost work often opens into adjacent tracks, and the sequencing is usually:

  1. Cost optimization — immediate, measurable, builds trust.
  2. Security hardening — the same inventory reveals open groups, missing MFA, unencrypted volumes. Your best-practice advisor reported these in stage 1.
  3. Availability and scalability — with waste removed and the estate understood, capacity work rests on real data.

Cost comes first because it is measurable, quick and self-funding. It is also how you earn the access and credibility the later phases need.


Topic 7: What Good Looks Like

At the end you should have:

  • A tagged estate where every resource has an owner and a cost centre.
  • An inventory snapshot defining engagement scope.
  • A layer-ranked analysis showing where money actually goes.
  • An implementation log with per-change evidence and rollback paths.
  • A monitoring posture — anomaly detection configured, budgets alerting — so the next regression is caught in days rather than at quarter end.
  • A re-baselining date, because the work decays without one.

That last item is what distinguishes a cost programme from a cost project. Baselines shift; without a scheduled review the estate drifts back.


Try it yourself: Run phases 1 and 2 fully on an account before implementing anything. Most people discover their assumed top cost layer is not the actual top cost layer, and that discovery is worth more than the first three optimizations.

Common mistake: Reporting the headline percentage without the per-resource evidence behind it. “We cut the bill 40%” invites the question “compared to what, and what else changed?” — and without per-day, per-resource attribution there is no answer, which turns a genuine result into an unverifiable claim.