Cloud cost work fails in a predictable way: someone finds an oversized instance, resizes it, reports a win, and stops. Three months later the bill is higher than when they started, because nobody looked at the other four layers where the money actually was.
Topic 1: Cost Is an Engineering Problem
The discipline goes by several names — FinOps, cost engineering, cloud financial management — and the framing that matters is this: cloud spend is a system property, like latency or availability. It responds to architecture, and it degrades without attention.
Unmanaged infrastructure compounds. Every environment someone spins up and forgets, every log group with no retention, every volume detached from a terminated instance keeps billing indefinitely. None of it fails loudly.
A real engagement shape, drawn from a healthcare platform’s estate:
| Environment | Monthly cost |
|---|---|
| Production | ~$45,000 |
| Non-production (dev + staging) | ~$53,000 |
| Total | ~$98,000 |
Read that table again. Non-production cost more than production. That single fact tells you almost everything before you open a single console: dev and staging were running 24×7 with no scheduling, and were almost certainly overprovisioned, because dev workloads are bursty and low-traffic by nature.
That is what cost engineering looks like — reading a bill as a diagnostic signal rather than a number to reduce.
Topic 2: The Layer Model
The framework that prevents tunnel vision. Work through the estate layer by layer, in order of cost impact:
Layer 1: COMPUTE
├── Server: virtual machines, Kubernetes nodes, committed compute
└── Serverless: functions, container services
Layer 2: NETWORK
├── NAT gateways
├── Inter-region and cross-AZ data transfer
├── Load balancers
└── Idle public IPs
Layer 3: STORAGE / DATA
├── Block volumes, snapshots, volume types
└── Object storage: lifecycle, tiering, archival
Layer 4: OBSERVABILITY
└── Logs, metrics, traces, retention
Layer 5: OTHER TOOLING
├── Managed databases and caches
├── Security tooling
└── Registries and artifact stores
Why a layer model at all? The same reason you troubleshoot a network by layer: it guarantees nothing is missed. Without it, cost work becomes whatever the engineer already knows how to fix — which is usually compute, because compute is visible.
Ordering rationale: start with the layer consuming the most cost, or default to compute since it is usually the largest driver, and work down. Both are defensible. What is not defensible is starting with the layer you find most interesting.
Topic 3: The Layer Order Is Not Universal
Compute is usually the biggest driver. Usually is not always, and assuming it wastes the whole engagement.
A fintech platform’s 90-day billing breakdown, grouped by billing line item:
| Component | 90-day cost | Share |
|---|---|---|
| Log storage | ~$15,000 | 44% |
| Virtual machines | ~$7,500 | 21% |
| Monitoring | ~$3,800 | 11% |
| Managed SQL | ~$1,500 | — |
| Managed cache | ~$1,400 | — |
That estate was not compute-heavy. Nearly half the avoidable spend sat in observability — the layer most engineers check last, if at all. Right-sizing every instance there would have addressed a fifth of the problem while the actual driver kept growing.
The lesson generalises: always let the billing data pick the layer. Group last month’s spend by service before forming any hypothesis.
Topic 4: The Three Categories
Every resource you inventory falls into one of three buckets:
| Category | Meaning | Action |
|---|---|---|
| Overprovisioned | Utilisation well below capacity | Downsize, or change architecture |
| Properly utilised | Capacity matches demand | Leave the size; optimise architecture only |
| Underprovisioned | Running hot, at risk | Upsize |
That last row surprises people during a cost engagement. If you find an underprovisioned resource, you increase its size and its cost. You are engineering the estate, not minimising a number — and an outage caused by a cost programme costs more than the programme saved.
Topic 5: What Triggers an Engagement
Cost work rarely starts because someone decided to be efficient. Four recognisable triggers, and each tells you something about what you will find:
| Trigger | What it usually means |
|---|---|
| A sudden bill spike | A misconfiguration, a runaway job, or a log level left at debug. Look for a change, not a trend |
| Consistently high bill against low utilisation — CPU under 20%, memory under 40% | Systemic overprovisioning. The whole estate needs right-sizing, not one instance |
| Recently migrated from on-premises or from a monolith | Instances sized like the physical servers they replaced. Almost always heavily overprovisioned |
| Recent lift-and-shift | Same problem, more acute — nothing was re-architected, so nothing is cloud-shaped |
The last two matter because they tell you the overprovisioning is structural rather than accidental. Somebody sized a virtual machine to match a physical one, and nothing since has revisited it. That is a large, safe, systematic saving — and it is why post-migration estates are the most rewarding cost engagements.
Topic 6: What Cost Work Is Not
It is not a one-time project. Baselines shift as traffic grows, features ship and seasons turn. A 90-day baseline taken once and never revisited becomes wrong quietly.
It is not free of risk. Every change — resizing, migrating architecture, moving to interruptible capacity — carries a blast radius. The discipline is applying them in an order where each step is individually reversible.
It is not the DevOps team’s decision alone. Consolidating availability zones to cut cross-zone transfer cost changes your failure domain. Scheduling a shutdown assumes nobody works those hours. Both need explicit business sign-off, and both are trivially easy to do without asking.
And it is not achieved without attribution. If you cannot say which team owns which spend, you cannot ask anyone to change behaviour — which is why the next lesson is about tagging before it is about saving.
Try it yourself: Group last month’s bill by service and calculate what percentage sits in each of the five layers. Most people are surprised by at least one number, and that surprise is the highest-value output of the exercise.
Common mistake: Starting with the resource you already know is oversized. It feels productive and it anchors the entire engagement to one layer. Inventory first, rank by cost, then choose — even when the ranking sends you somewhere unfamiliar.