The Runaway Bill
A Cloud Bill Growing Faster Than Traffic, With Nobody Able to Say Whose It Is
The situation you’re stepping into
A platform spanning two clouds. Managed Kubernetes in both, a spread of managed databases and caches, object storage, and a hosted observability platform taking logs and metrics from everything.
Monthly spend has crossed budget and is still rising. Traffic is flat. No incident has occurred, no alert has fired, and when finance asks which teams are responsible for the increase, nobody can answer — not because the teams are evasive, but because the data to answer it does not exist.
This is the one entry in this collection that is not an availability incident. It is included because the failure mode is the same shape as the others: a system degrading steadily while every health signal reports fine. The difference is that the damage is financial and compounding rather than sudden, which makes it easier to ignore for longer.
What the team observed
- Month-over-month spend up materially, with request volume flat.
- No single line item obviously responsible — the increase is spread.
- Roughly a third of resources carry no owner tag of any kind.
- Non-production environments are running continuously, including overnight and at weekends.
- The Kubernetes clusters show low average CPU utilisation and are nonetheless scaling out.
- The observability platform’s ingestion line has become one of the largest single items on the bill.
- Nobody has an inventory. The only view anyone has is the provider console, one service at a time.
?
Decision Point 1 You have a rising bill and no attribution. What do you do FIRST — and why is finding an oversized instance and resizing it the wrong opening move, even though it would show an immediate saving?
Ask what you would be able to prove at the end of the engagement, and to whom.
Commit to your answer, then reveal the responder’s move
→
You have a rising bill and no attribution. What do you do FIRST — and why is finding an oversized instance and resizing it the wrong opening move, even though it would show an immediate saving?
Ask what you would be able to prove at the end of the engagement, and to whom.
Commit to your answer, then reveal the responder’s move →Establish attribution and a baseline before changing anything. Specifically: a tag audit, a full resource inventory, and the provider’s native cost tooling switched on.
Resizing something first is tempting because it produces a number quickly. It is the wrong opening move for three reasons:
You cannot prove the saving. Bills move for reasons unrelated to your work — traffic growth, a new service, a pricing change. Without a baseline and per-resource attribution, “the bill went down 12%” is a claim, not evidence. The first question in the review will be “compared to what, and what else changed?” and you need an answer.
You cannot change behaviour. The bill is rising because teams create resources without a feedback loop. Nothing you resize prevents the next untagged, oversized deployment. Attribution is what converts a central cost problem into a distributed one that the people creating the spend can actually see.
You will attack the wrong layer. Compute is usually the largest driver, so it is where instinct points. On this platform the observability line turned out to rival compute. Reaching for instance sizes first would have addressed a fraction of the problem while the actual driver kept growing.
The opening sequence is therefore:
# 1. Tag coverage — what percentage of resources have a complete tag set?
# 2. Full inventory — every resource, type, size, state, tags, utilisation, cost
# 3. Enable native cost tooling (explorer, right-sizing recommender, advisor, anomaly detection)
# 4. Start a 90-day utilisation baseline running in the background
Step 3 matters more than it looks: the recommenders are historical analysers. Enabled on the day you need their output, they have nothing to say. Enabled on day one, they are useful by week four at no cost.
?
Decision Point 2 A third of resources are untagged. Why is tagging the single highest-value action here, and what does a tag schema actually need to contain to be useful?
Think about what each tag unlocks that is impossible without it — not what it documents.
Commit to your answer, then reveal the responder’s move
→
A third of resources are untagged. Why is tagging the single highest-value action here, and what does a tag schema actually need to contain to be useful?
Think about what each tag unlocks that is impossible without it — not what it documents.
Commit to your answer, then reveal the responder’s move →Because every downstream capability depends on it, and none of them degrade gracefully in its absence — they simply become impossible.
A workable schema is four tags, mandatory on every resource regardless of how it was created:
| Tag | Unlocks |
|---|---|
Name | Human identification |
Owner | Safe deletion — you can ask before removing something |
Environment | The automation filter. Every cost automation targets this |
CostCentre | Per-team reporting and chargeback |
CostCentre is what makes the engagement stick. Without it there is no team-level report, so no team sees the consequence of its own provisioning, so the bill resumes climbing the moment the engagement ends. With it, spend becomes a number each team owns and can be asked about.
Environment is the safety mechanism, not merely a label. Every piece of cost automation you write — shutdown schedules especially — targets it. Production never matches the filter, so it is never in scope. That is the difference between an automation you can run unattended and one somebody has to supervise.
Enforce it in three places, because hand-applied tags decay:
- Policy at provision time — reject creation without the required set. The only mechanism that prevents rather than reports.
- Infrastructure-as-code defaults — a shared tag map applied to every resource, so tagging is not something anyone has to remember.
- Scheduled audit — report untagged resources weekly to whoever created them.
A useful transitional rule: untagged non-production resources get stopped after a grace period. It converts a nagging problem into a self-solving one inside a fortnight.
?
Decision Point 3 Your inventory is complete and the baseline is running. The clusters show low average CPU and are still scaling out. What is actually happening, and why will resizing nodes not fix it?
The scheduler places pods on what they ask for, not on what they use.
Commit to your answer, then reveal the responder’s move
→
Your inventory is complete and the baseline is running. The clusters show low average CPU and are still scaling out. What is actually happening, and why will resizing nodes not fix it?
The scheduler places pods on what they ask for, not on what they use.
Commit to your answer, then reveal the responder’s move →The pods are over-requesting, and node autoscaling is faithfully providing capacity to satisfy requests that bear no relationship to consumption.
This is the single largest source of hidden waste in a Kubernetes estate, and it is invisible to every node-level tool:
- The scheduler places pods on
requests, not usage. A pod requesting 4 CPU and using 0.3 consumes 4 CPU of schedulable capacity. - The node autoscaler provisions nodes large enough to satisfy those requests. It is working perfectly and has no visibility into whether the request was accurate.
- Average node CPU therefore stays low while the cluster keeps growing — the exact signature on this platform.
Resizing nodes does not touch the cause. Smaller nodes satisfying the same inflated requests means more nodes, and possibly a worse bin-packing outcome.
The fix is layered, in this order:
- Right-size requests against observed usage per workload. This is where the money is.
- Then improve node provisioning — demand-based provisioning that picks the cheapest instance fitting the (now honest) requests, with consolidation to reclaim idle nodes off-peak.
- Then evaluate cheaper capacity: interruptible instances for tolerant workloads, and alternative CPU architectures where the dependency scan comes back clean.
The distinction to hold onto: an autoscaler optimises provisioning against a request. A cost-observability tool tells you whether the request itself was right. Different data, complementary jobs — which is the answer to “we already have an autoscaler, why do we need a cost tool”.
?
Decision Point 4 The observability platform's ingestion line rivals compute. How do you reduce it without losing the data you debug with — and which lever do you pull first?
Cost here is roughly volume multiplied by retention. One of those two is much safer to change.
Commit to your answer, then reveal the responder’s move
→
The observability platform's ingestion line rivals compute. How do you reduce it without losing the data you debug with — and which lever do you pull first?
Cost here is roughly volume multiplied by retention. One of those two is much safer to change.
Commit to your answer, then reveal the responder’s move →Retention first, volume second — because halving retention halves that component’s storage cost without losing any recent data, and recent data is what you actually debug with.
Cost here is approximately volume × retention, two independent levers:
Retention is faster, safer and reversible going forward. The question to ask each owning team is not “can we cut this to 14 days?” but “how far back have you actually needed to look in the last year?” Retention is almost always set to a default nobody chose, and the honest requirement is usually far shorter than the configured value.
Volume requires deciding what not to emit. Debug-level logging in production generates 10–100× the volume of info or warn, and is very often left on after an investigation. Health-check and load-balancer access logs are frequently the majority of volume and near-zero diagnostic value.
Then the adjacent items that cost money quietly: unused dashboards, custom metric indexes nobody queries, and retention on metrics as distinct from logs.
Two things must not be cut:
- Compliance-mandated retention. Identify which streams are fixed first, so you do not present a saving you then have to withdraw.
- Enough metric history for capacity work. Right-sizing needs 90 days of utilisation. Cutting metric retention to 30 days undermines the compute work you are doing in parallel — a genuine tension between two parts of the same engagement, and one to decide deliberately.
Filtering and restoration are not symmetric operations. If you write an exclusion rule and later need that data, there is no restoration — it was never stored. A “rollback” here only stops excluding going forward. Say this explicitly to whoever signs off, because “we can roll it back” means something very different than it does for a config change.
?
Decision Point 5 Storage and networking are the layers nobody has looked at. What is safe to delete immediately, and where is the trap that can turn a cost engagement into an incident review?
Some of this is pure waste with no trade-off at all. One item on the list changes your failure domain.
Commit to your answer, then reveal the responder’s move
→
Storage and networking are the layers nobody has looked at. What is safe to delete immediately, and where is the trap that can turn a cost engagement into an incident review?
Some of this is pure waste with no trade-off at all. One item on the list changes your failure domain.
Commit to your answer, then reveal the responder’s move →Work this layer waste-first, because the early items carry no risk and buy credibility for the later ones.
Pure waste — no trade-off, no conversation required:
- Unattached block volumes. Terminating an instance does not necessarily remove its volumes; detaching certainly does not. They bill at full price indefinitely.
- Snapshot sprawl. Cheap individually, unbounded collectively. Automated daily snapshots with no expiry grow forever.
- Idle public addresses. An allocated address not attached to a running resource bills continuously.
- Unused load balancers. Each carries an hourly charge whether or not anything is behind it.
- Stopped instances’ storage. A stopped instance costs nothing for compute and continues to charge for every attached volume.
Low risk, clear benefit:
- Volume type migration to a newer generation — typically cheaper per GB and better baseline throughput. One of the few changes that is unambiguously better on both axes.
- Object lifecycle policies to transition ageing data to cheaper tiers. Check retrieval cost and minimum storage duration before being aggressive — data you might need during an incident does not belong in deep archive.
- Private endpoints for provider-service traffic, so cluster traffic to object storage stops traversing a managed gateway that bills per GB. Frequently the largest single networking saving available.
The trap — and it is a real one:
Consolidating availability zones reduces cross-zone transfer cost and reduces your fault tolerance. Same for collapsing redundant load balancers or simplifying peering topology. These change your failure domain.
Never make an availability-topology change for cost reasons without explicit, written business sign-off. This is the one category in the entire engagement that can convert a cost saving into an outage, and it is the reason architectural items sit at the bottom of the plan rather than the top.
?
Decision Point 6 Everything above is done and the bill is falling. A colleague proposes locking in a multi-year commitment discount for the deepest possible saving. Why is the timing of that decision the most consequential judgement in the engagement?
What is the commitment sized against, and what have you spent the last two months doing to that number?
Commit to your answer, then reveal the responder’s move
→
Everything above is done and the bill is falling. A colleague proposes locking in a multi-year commitment discount for the deepest possible saving. Why is the timing of that decision the most consequential judgement in the engagement?
What is the commitment sized against, and what have you spent the last two months doing to that number?
Commit to your answer, then reveal the responder’s move →Because a commitment is the only decision here you cannot undo, and every optimization you have just performed has changed the number it should be sized against.
Commitment discounts trade flexibility for a discount over one to three years. The rules that matter:
- They apply to on-demand capacity, not interruptible capacity. Decide the Spot strategy first, then commit against what remains. Committing against spend that Spot is about to remove buys nothing.
- Commit to the baseline floor, never the peak. Chart daily spend across 90 days and draw a line at the minimum. That floor is the defensible commitment level. Under-committing leaves some saving unclaimed; over-committing is a loss carried for years.
- There is no exit. They generally cannot be cancelled and have no secondary market.
- Plan type matters more than amount. The flexible compute plan covers serverless and managed container compute; the family-specific plan does not. Picking the wrong type is a more common and more expensive mistake than picking the wrong amount.
The sequencing rule, stated plainly: buy the commitment last. Right-sizing, architecture migration, scheduling and Spot adoption all reduce on-demand spend. Buy early and you lock in the pre-optimization baseline for up to three years — paying for capacity your own successful work has just made unnecessary.
This is the item most likely to be pushed forward, because it shows the largest headline saving and requires no engineering. Resisting that is the judgement being tested.
Root cause
There was no attribution, so there was no feedback loop. Teams created resources without visibility into their cost and without anyone able to associate spend with a decision. The bill therefore grew as a function of headcount and deployment velocity rather than traffic — which is exactly why it kept climbing while request volume stayed flat.
The specific drivers, in the order the data ranked them:
- Over-requested Kubernetes workloads driving node provisioning far beyond real utilisation.
- Observability ingestion and retention left at defaults nobody had ever chosen.
- Non-production environments running continuously, at roughly three times the hours anyone used them.
- Accumulated storage and networking waste — unattached volumes, snapshots without expiry, idle addresses, orphaned load balancers — none of it owned by anyone.
- Untagged resources, which made all of the above invisible until an inventory was built.
No component failed. Every system did exactly what it was configured to do.
Resolution and prevention
Implementation ran risk-ascending, so early zero-risk wins funded credibility for the harder conversations:
| Phase | Work | Risk |
|---|---|---|
| 1 | Tag audit and remediation; full inventory; enable native tooling; start 90-day baseline | None |
| 2 | Delete unattached volumes and snapshots, release idle addresses, remove orphaned load balancers | None |
| 3 | Non-production shutdown scheduling, tag-filtered | Low |
| 4 | Volume type migration; object lifecycle policies | Low |
| 5 | Observability retention and volume reduction | Low–medium |
| 6 | Workload request right-sizing; node provisioning improvements | Medium |
| 7 | Interruptible capacity and architecture migration where the dependency scan was clean | Medium |
| 8 | Commitment purchase against the post-optimization baseline | Last |
Everything was codified in infrastructure-as-code rather than applied by hand — tag schema, shutdown schedules, lifecycle policies, node pool configuration, retention settings. Console changes do not survive the next deployment, cannot be reviewed, and leave no audit trail.
Verification per change: before-and-after cost from the cost explorer grouped by day, plus a performance check confirming no regression. A saving without performance evidence is an unverified claim; a monthly bill total attributes nothing.
Prevention:
- Tag enforcement at provision time so the attribution gap cannot reopen.
- Per-team chargeback reporting so each team sees its own spend as a routine number rather than an annual surprise.
- Budget alerts with threshold notifications, and anomaly detection per service — the only forward-looking control in the whole programme.
- A scheduled re-baselining date. Baselines shift as traffic and features change. Without one, the estate drifts back and the work decays.
The pillar to state explicitly to any stakeholder: cost is one of three, alongside performance and reliability. A resource that is genuinely under-provisioned should get bigger during a cost engagement — costing more, and being the correct call. A programme that causes an incident has not saved anything.
Telling this story to a recruiter
Situation. Cloud spend across two providers had breached budget and was compounding month over month while traffic stayed flat. Roughly a third of resources carried no owner tag, no inventory existed, and finance could not attribute spend to any team — so there was no feedback loop between the people creating cost and the cost being created.
Task. Reduce spend materially across both clouds without degrading performance or reliability, and leave behind an attribution model so the problem could not silently reopen.
Action. Started with foundations rather than quick wins: audited tag coverage, built a full resource inventory that also served as the engagement’s scope boundary, enabled native cost tooling on day one so recommenders had history, and started a 90-day utilisation baseline. Ranked spend by layer from the billing data rather than assuming compute dominated — which surfaced observability ingestion as rivalling compute, a layer that would otherwise have been checked last. Right-sized Kubernetes workload requests through profiling before touching node provisioning, since the scheduler places pods on requests rather than usage and node-level tools cannot see that gap. Improved autoscaling policy, and evaluated interruptible capacity and alternative CPU architectures for workloads where the compatibility assessment was clean. Cut observability cost by auditing log ingestion volume and pruning unused dashboards, indexes and retention policies, pulling the retention lever before the volume lever because it is reversible going forward. Automated cleanup of unattached volumes and snapshots, applied object storage lifecycle policies, and eliminated networking waste — orphaned load balancers, idle public addresses, and inefficient routing. Enforced tagging and cost-allocation labels through Terraform across every environment, and codified every change in infrastructure-as-code rather than the console. Sequenced the work risk-ascending and deliberately left the commitment-discount purchase until last, so it was sized against the optimized baseline rather than the original one.
Result. 25–40% reduction in monthly cloud spend across both providers, with observability cost down approximately 20% on its own, and measurably improved infrastructure utilisation. Every change was evidenced with per-resource, per-day cost data alongside a performance check, so the savings were attributable rather than merely coincident with the work. Real-time FinOps visibility and team-level chargeback reporting were established, converting a central cost problem into one the teams generating the spend could see and own.
What this demonstrates. Treating cost as an engineering property rather than a finance report; sequencing work by risk so that the irreversible decision comes last; and holding the line that a saving which degrades reliability is not a saving.
Interview deep-dive: the full case study
Why attribution is the first move and not the last. A cost programme that reduces spend without establishing ownership produces a one-time dip followed by a resumption of the original growth curve, because nothing about the behaviour that created the spend has changed. Tagging is what converts a central problem into a distributed one: each team sees its own number, is asked about it routinely, and adjusts. Every other intervention in the engagement is a one-off; attribution is the only structural change. It is also the least glamorous, which is why it gets skipped.
Why the billing data must choose the layer. The instinct is to start with compute, and compute usually is the largest driver — but “usually” is doing a lot of work in that sentence. On this platform the observability line rivalled compute, and an engineer following instinct would have spent weeks right-sizing instances while the actual driver compounded. The discipline is to group last month’s spend by service before forming any hypothesis, and to follow the ranking even when it sends you somewhere unfamiliar.
Why requests, not nodes, are the Kubernetes cost story. The scheduler allocates on requests. A pod asking for four CPUs and using a third of one consumes four CPUs of schedulable capacity, and the node autoscaler dutifully provisions hardware to satisfy it. Average utilisation therefore stays low while the cluster grows — a signature that looks like an autoscaler misconfiguration and is not. No node-level tool can detect it, because the autoscaler is behaving correctly against the information it has. This is the cleanest example of why an autoscaler and a cost-observability tool are complementary rather than redundant: one optimises provisioning against a request, the other tells you whether the request was honest.
Why retention before volume. Both reduce cost and they differ sharply in risk. Reducing retention loses old data on a schedule you can extend again; it costs one conversation and is reversible going forward. Reducing volume means deciding what never to emit, involves more stakeholders, and — critically — is not reversible: a filtered log was never stored, so there is nothing to restore. Teams routinely promise a “rollback” for log filtering without realising it only stops future exclusion. Being precise about that distinction with whoever signs off is what stops a cost saving becoming a debugging outage later.
Why the commitment goes last, and why that is the hardest thing to hold. Every prior optimization reduces on-demand spend. A commitment bought first is sized against the pre-optimization baseline and locks it in for up to three years — you end up paying for capacity your own work made unnecessary, with no exit and no secondary market. It is the only irreversible decision in the programme, and it is simultaneously the one with the largest headline number and the least engineering effort, which is exactly why it gets pulled forward. Being able to explain the sequencing argument clearly is a strong signal in an interview, because it demonstrates you understand the interaction between the levers rather than knowing them as a list.
Why cost work needs an explicit reliability position. The framing worth stating up front is that cost sits alongside performance and reliability, and is never optimised at their expense. The practical consequence is that some resources will not be resized and one or two will get larger — an under-provisioned resource found during a cost engagement should be increased, costing more, because that is the correct engineering call. Saying this to stakeholders at the start sets the expectation that the deliverable is a well-engineered estate rather than a minimised number, and it is what gives you the standing to refuse the availability-topology change that would have saved money and reduced the failure domain.
What to say when asked “how do you know the saving was real?” Per-resource, per-day cost evidence from the cost explorer, paired with a performance check for each change. A monthly bill total proves nothing — bills move for traffic growth, new services and pricing changes, none of which are your work. The inventory snapshot taken at the start also matters here: it defines what was in scope, so resources provisioned afterwards are demonstrably outside the engagement. That artifact has settled more disputes than any optimization in it.