Observability Cost — Logs, Metrics & Retention

The layer that is 44% of the bill on some estates: separating retention from volume, distinguishing audit logs you cannot touch from product logs you can, and cutting spend without losing the data you debug with.

advanced 20 min lesson hands-on task included

Observability is the layer engineers check last and the one that most often tops the bill. On one fintech estate, log storage alone was 44% of total cloud spend — more than double the compute line — in an environment nobody would have described as logging-heavy.


Topic 1: Why It Grows Unnoticed

Four properties combine badly:

  • Retention frequently defaults to indefinite. Nothing expires unless you configure expiry.
  • Volume scales with deployment count, not with traffic. More services means more streams, whether or not anyone reads them.
  • A log level is a one-line change with a 10–100× cost multiplier. Debug logging left on in production after an investigation is the classic version.
  • Nobody owns it. Logs are produced by every team and paid for centrally.

The result: a line item that grows monotonically and is never reviewed, because reviewing it requires someone to decide what can be deleted.


Topic 2: Separate Volume from Retention

Two independent levers, and conflating them produces bad decisions.

  • Volume — how much you ingest. Controlled by log level, sampling, and what you choose to emit at all.
  • Retention — how long you keep it. Controlled by policy.
Cost ≈ volume × retention — two independent levers volume retention current bill after both levers 30d → 14d halves storage drop debug · info · stdout · noisy namespaces Filtering and restoration are not symmetric operations. A "rollback" only stops excluding data going forward. Anything already dropped is gone permanently.
Two independent levers multiply. Retention is the faster and safer one to pull first — and the red band is the asymmetry that catches teams out when they promise a rollback.

Cost is roughly the product of the two. Halving retention halves storage cost without losing any recent data — which makes it almost always the right first move, because recent data is what you actually debug with.

Halving volume also halves ingestion cost, but requires deciding what not to emit, which is a longer conversation with more stakeholders.

Start with retention. It is faster, safer and reversible going forward.


Topic 3: Not All Log Streams Are Equal

A pattern that generalises across providers. A real estate had two log destinations:

DestinationContentsRetentionNegotiable?
DefaultApplication and infrastructure logs30 daysYes
Required/auditEverything in default, plus third-party integration logs for compliance400 daysNo — fixed baseline

Two things matter here.

First, the audit destination is a superset. It contains everything the default one does plus more. Understanding that relationship before proposing changes prevents recommending a reduction that breaks a compliance obligation.

Second, the 400-day retention was a fixed compliance baseline that could not be reduced at all. No amount of cost pressure changes it. Identifying which streams are immovable first stops you presenting a saving you then have to withdraw.

The question to ask the owning team: how far back have you actually needed to look in the last year? On that estate the answer for product logs was 7–14 days against a configured 30. Reducing 30 days to 14 was projected to cut that destination’s storage cost by roughly half.

That question is the whole technique. Retention is almost always set to a default nobody chose, and the real requirement is usually far shorter than the configured value.


Topic 4: The Levers, Ranked

1 — Set retention deliberately. Every log group, every bucket, an explicit number derived from a stated need. Groups with no retention policy are the highest-value target in the whole layer.

2 — Fix production log levels. Debug in production generates 10–100× the volume of info or warn. Audit what each service emits at, and check for the one somebody switched during an incident and never switched back.

3 — Tier old logs to cheap storage. Export beyond the hot window to object storage and let a lifecycle policy transition it to archive. Retains the data for compliance at a fraction of the platform’s own storage price. Verify retrieval time is acceptable for whatever you might need it for.

4 — Prune metrics and alarms. Custom metrics bill per metric. Dashboards nobody opens and alarms nobody acts on still cost money every month.

5 — Sample high-volume, low-value streams. Health check and load balancer access logs are often 90% of volume and near-zero diagnostic value. Sample or drop them at the source.

6 — Evaluate self-hosting. A self-managed log stack can be materially cheaper at high volume. Weigh honestly: you are trading a bill for operational burden, and a self-hosted logging stack you cannot keep running during an incident is worse than an expensive one that stays up.


Topic 5: Filtering Is Not Reversible

The single most important operational caveat in this lesson, and the one most likely to be promised carelessly.

If you write a filter that drops a category of logs and later decide you need it, there is no restoration. The data was never stored. A “rollback script” in this context is a forward-facing mechanism: it stops the exclusion so you begin capturing that category again from that moment. It does not recover anything that was already dropped.

Filtering and restoration are not symmetric operations. Say this explicitly to whoever signs off, because “we can roll it back” means something quite different here than it does for a configuration change, and the gap between those two meanings is discovered at the worst possible time.

Practical consequence: be more conservative with exclusion filters than with retention. Shortening retention loses old data on a known schedule you can extend again. An exclusion filter loses data you never had the option to keep.

A realistic categorisation exercise:

Break the stream into categories before proposing anything — application, infrastructure, debug, info, stdout, per-namespace — and take that breakdown to the owning team. On one estate that conversation produced agreement to drop debug, info and stdout entirely, and to exclude platform-namespace logs from the cluster, none of which were part of any audit requirement.

One honest constraint worth naming: the cleanest fix is usually to stop emitting the logs at source, in the application. On that same estate the team proposed exactly that and the client declined — so the entire reduction had to be implemented at the infrastructure layer through exclusion rules and retention. The technically better answer is not always the one you are permitted to implement, and recognising that early saves a fortnight of arguing for it.


Topic 6: What Not to Cut

Cost work on this layer has a failure mode worse than overspending: deleting the data you need at the exact moment you need it.

  • Keep the recent hot window generous. The 7–14 days you debug with is the cheapest part of the dataset because it is the smallest.
  • Never reduce compliance-mandated retention. Confirm what is mandated rather than assuming — and confirm it with whoever owns the obligation, not with another engineer.
  • Do not sample error and audit paths. Sample the health checks, never the exceptions.
  • Keep enough history for capacity work. The baselining from stage 1 needs 90 days of metrics. Cutting metric retention to 30 days undermines the right-sizing you did two lessons ago.

That last point is a genuine tension between two parts of this module, and it is worth deciding explicitly rather than discovering later.


Topic 6: Building the Case

Observability cost reduction needs agreement from teams who did not create the budget problem and will feel the loss. Make it concrete:

  • Show the share. “Logging is 44% of the bill” reframes the conversation instantly.
  • Ask for the real requirement, do not propose a number. “How far back have you needed to look?” gets a better answer than “can we cut this to 14 days?”
  • Model before changing. State the projected saving and the projected loss of coverage together.
  • Separate the immovable. Naming what you are not touching makes the rest credible.

One honest caution on the arithmetic. Live cost estimates in this layer are easy to get wrong — volume figures quoted per day, per month and per billing period get mixed up, and the intermediate numbers stop reconciling. The direction is reliable: halving retention roughly halves that component’s storage cost. Verify the specific figures against actual billing data before presenting them.


Try it yourself: List every log group with its retention setting and sort by ingestion volume. The groups at the top with no retention policy set are your entire first week of work in this layer.

Common mistake: Cutting retention across the board with a single policy. Audit streams have fixed compliance requirements, security streams may need longer than product logs, and one global number will either break a compliance obligation or save far less than a targeted policy would.