The Native Cost Toolchain

The four provider-native tools to enable before any optimization work, what each is actually good for, how they differ across clouds, and using cost explorers to prove a change worked.

intermediate 18 min lesson hands-on task included

Every major cloud ships tooling that does the first pass of cost analysis for you. It is consistently underused — partly because it needs enabling before it collects anything, and partly because engineers assume tools this easy cannot be saying anything they do not already know.


Topic 1: The Four Tools

Enable all four before optimization work begins. Several need a day or more of data collection before they show anything useful.

Tool categoryWhat it doesPriority
Cost explorerVisualises spend by service, resource, tag and time rangeFirst
Right-sizing recommenderRecommends instance sizes from observed utilisationSecond
Best-practice advisorIdle resources, security gaps, quota warnings, reliability findingsThird
Anomaly detectionAlerts on unexpected spend spikes per service or accountFourth
Enable all four on day one — several need a day of history before they say anything 1 · EXPLORER where to look group by day to prove 2 · RECOMMENDER what to change needs ~24h+ of data 3 · ADVISOR what you missed idle · security · quota 4 · ANOMALY the next surprise only forward-looking one Where they stop They see resources, not architecture — and rarely reach inside a cluster to notice a pod requesting six times what it uses. Ordering reflects dependency, not importance: the explorer tells you where to look before the recommender is worth reading. Cross-cloud, one provider's advisor is free and another gates it behind paid support.
Enable all four before you decide anything. Three of them are historical analysers — switched on the morning you need output, they have nothing useful to say that afternoon.

The ordering reflects dependency, not importance. The explorer tells you where to look; the recommender tells you what to change; the advisor catches what you did not think to check; anomaly detection stops the next surprise.


Topic 2: Cost Explorer

The workhorse. Its most valuable use is not exploration but proof.

After any change, the verification loop:

  1. Filter to the resource type, then to the specific resource ID.
  2. Filter by the new instance type or configuration.
  3. Group by day — this is the step people miss, and without it the change is invisible inside a monthly total.
  4. Compare the day before the change against the day after.
  5. Export or screenshot for the change record.

A concrete result from that loop: an instance migrating between CPU vendors moved from roughly $20/day to $16–17/day — about $100–130 per month, for one instance, from a change with a two-minute maintenance window.

Costs typically settle within 24 hours of a change. Verify then, not at month end when the signal is buried.


Topic 3: The Right-Sizing Recommender

Analyses collected utilisation and recommends a size, with a projected saving attached.

  • Must be enabled, and typically needs up to 24 hours before showing data.
  • Recommendations improve as history accumulates — treat week-one output as provisional.
  • The equivalent tool exists on every major cloud under a different name; the concept is identical.

Its real value is removing guesswork and, more importantly, removing argument. “This instance should be smaller” is an opinion. “The provider’s own recommender projects a 60% saving from this change based on 90 days of utilisation” is a proposal.

Do not implement recommendations mechanically. The recommender sees utilisation, not context — it does not know about the quarterly batch job, the failover capacity you deliberately hold, or next month’s launch.


Topic 4: The Best-Practice Advisor

Broader than cost, which is what makes it useful. Typical findings across categories:

CategoryExample findings
CostIdle instances, unattached volumes, underutilised databases
SecurityAccounts without MFA, security groups open to the world on admin ports
PerformanceMissing autoscaling, misconfigured health checks
LimitsResources approaching a service quota
ReliabilityMissing disruption budgets, short backup retention, absent maintenance windows

The cost findings here are the fastest money in the whole discipline — unattached volumes and idle public IPs are pure waste with no performance consideration and no business conversation required. One estate cleared roughly $232/month from unattached block volumes alone.

A cross-cloud difference worth knowing: one major provider’s equivalent tool is entirely free, while another gates the full recommendation set behind a paid support tier. If you work across clouds, do not assume parity.


Topic 5: Anomaly Detection

The only tool here that is forward-looking. Everything else describes what already happened; anomaly detection tells you about the spend spike while it is still happening.

  • Configure per service and per account, with a threshold that reflects normal variance.
  • Route alerts somewhere a human reads — a cost alert in an unmonitored inbox is not a control.
  • Tune the threshold. A detector that fires on every deployment gets muted, and then it is worse than nothing.

This is what catches the misconfigured job that spins up instances in a loop, or the log level someone set to debug in production on a Friday.


Topic 6: Where Native Tools Stop

Being clear about the ceiling saves you from expecting too much:

  • They see resources, not architecture. No tool will tell you the workload should not run in that region, or that two services could share a cluster.
  • They cannot see business context. Deliberately-held failover capacity looks identical to waste.
  • They rarely reach inside Kubernetes. A cluster appears as nodes; the recommender cannot see that a pod requests six times the memory it uses. That gap is exactly what the Kubernetes-specific tooling in stage 3 exists to close.
  • Recommendations are per-resource. They do not compose. Twenty recommendations, each individually sound, can add up to an architecture nobody designed.

Try it yourself: Export every recommendation with its projected saving, then rank the list twice — once by saving, once by implementation risk. The disagreement between those two orderings is your actual implementation plan.

Common mistake: Enabling the tools on the day you need their output. They are historical analysers — a recommender enabled this morning has nothing useful to say this afternoon. Enable them on day one of any engagement, before you have decided what to do.