“Tagging is the most important thing if you are talking about cost engineering in any cloud provider.”
Everything downstream depends on it. Without tags you cannot attribute spend to a team, automate anything safely, or answer the only question that matters in a cost review: who owns this, and do they still need it?
Topic 1: The Tag Schema
Four tags, mandatory on every resource regardless of how it was created — console, IaC, or CLI:
| Tag | Example | Purpose |
|---|---|---|
Name | platform-prod-node-1 | Human-readable identification |
Owner | platform-team | Accountability — who to ask before deleting |
Environment | prod / staging / dev | Cost split, and the key automation targets |
CostCentre | line-of-business-2 | Attribution to the team that pays |
Why the cost-centre tag matters most. In an organisation with multiple business units, each is accountable for its own spend. When a unit is audited and asked why it deviated from budget, the answer requires a per-team cost report. Without an attribution tag, that report cannot be produced at all — not with difficulty, but at all.
Why the environment tag matters second. It is the safety mechanism for every piece of cost automation you will write. A shutdown script targets Environment=dev. Production is never in scope, because it never matches the filter. That is the difference between an automation you can run unattended and one you have to babysit.
Topic 2: Enforcing Tags
Tags applied by hand decay. Three mechanisms that actually hold:
- Policy at provision time. Reject resource creation without the required tag set. This is the only approach that prevents the problem rather than reporting it.
- IaC defaults. A shared
common_tagsmap applied to every resource, so tagging is not a thing anyone has to remember. - Scheduled audit. Report untagged resources weekly and route them to whoever created them.
A useful transitional rule: untagged resources in non-production get stopped after a grace period. It converts a nagging problem into a self-solving one within about two weeks.
Topic 3: The Inventory
Before changing anything, build a point-in-time inventory of every resource: type, size, state, tags, utilisation, and current cost. One row per resource, grouped by category.
inventory/
├── compute/ # instances: name, zone, type, state, tags, attached IPs, disks
├── kubernetes/ # clusters, node pools, node labels and taints
├── storage/ # volumes, snapshots, buckets and their lifecycle rules
├── database/ # managed SQL, caches: tier, size, availability config
├── network/ # public IPs, load balancers, forwarding rules
└── observability/ # log buckets, retention, monitoring config
Pull each category to CSV with the provider CLI, then consolidate into a single workbook — one sheet per category. A short script using a dataframe library does this in a few lines and is worth writing once.
The inventory has a second purpose most people miss: it is an audit trail. If the estate grows after your engagement and someone observes that costs did not fall, the inventory snapshot defines exactly what was in scope. Anything provisioned afterwards is demonstrably not your work. That protection is worth the effort on its own.
Node labels and taints belong in the inventory. They determine which workloads land on which nodes, and therefore what you are paying for. A cluster inventory without them cannot support a node pool discussion later.
Topic 4: Baselining
You cannot right-size against a snapshot. You need a utilisation baseline.
Minimum window: 90 days. Longer for anything with seasonal variation — retail, healthcare reporting cycles, financial year-ends, education terms. A year of data is not excessive for a stable production workload.
Per resource, collect:
- CPU utilisation — average, P95 and P99, not just average
- Memory utilisation — usually requires an agent, and is the metric most often missing
- Disk IOPS and throughput, for storage-bound workloads
- Network throughput in and out
Average utilisation is the wrong number to size against. A service averaging 20% CPU with a P99 of 85% is not overprovisioned — it is correctly sized for its peak. Right-sizing on the average is how a cost programme causes an incident.
The window has honest limits:
Ninety days can still miss a seasonal spike or a traffic pattern introduced by a recent feature release. Baselining is continuous and evolving, not a one-time measurement. Re-baseline after any major release, and before any second round of right-sizing.
When the baseline is not yet trustworthy:
| Situation | Why |
|---|---|
| Within 30 days of provisioning | Not enough data |
| Rapidly growing traffic | Headroom assumptions expire too quickly |
| Mid-migration between providers | vCPU and memory performance ratios differ between platforms; right-sizing during a move compounds two variables |
| Just after a major feature release | Traffic pattern may have changed materially |
Topic 5: Sequencing
The order is not arbitrary:
- Tag — so everything after this is attributable.
- Inventory — so you know what exists and have your audit trail.
- Enable the native cost tools — so recommendations start accumulating data (some need 24 hours or more before they show anything).
- Baseline — 90 days minimum, running in the background while you do other work.
- Categorise — overprovisioned, properly utilised, underprovisioned.
- Plan per layer — expected savings in both percentage and currency, per service.
Step 3 is the one people leave until they need it. The recommendation engines are only as good as the history they have collected, so enabling them on day one costs nothing and makes week four far more productive.
Try it yourself: Calculate your tag coverage as a percentage of resources with a complete tag set. Under 80% means every cost report you produce is partly fiction, and fixing that is higher-value than any single resize.
Common mistake: Skipping the inventory because the billing console already shows costs by service. The console tells you what a service costs; it does not tell you which resources within it are idle, unattached, or owned by a team that dissolved last year. That resolution is where the savings actually live.