Compute is visible: an instance has a name, an owner and a purpose. Storage and network charges accumulate from things nobody decided to create — a volume left behind by a terminated instance, an address allocated for a migration two years ago, a bucket with no expiry.
Topic 1: Block Storage — Pure Waste First
Unattached volumes keep billing. Terminating an instance does not necessarily remove its volumes; detaching one certainly does not. They persist, at full price, indefinitely.
This is the cleanest money in cost engineering: no performance consideration, no business conversation, no risk beyond confirming the data is genuinely unwanted. One estate cleared roughly $232/month from unattached volumes alone.
The procedure: list every volume in an available (unattached) state, check its age and tags, snapshot anything uncertain, then delete.
Snapshot sprawl is the same problem one layer along. Snapshots are cheap individually and unbounded collectively — automated daily snapshots with no expiry policy grow forever. Set a retention rule and apply it retroactively.
Stopped instances still bill for their storage. A stopped instance costs nothing for compute and continues to charge for every attached volume. An estate full of “temporarily” stopped instances is paying storage rent on machines nobody intends to start.
Topic 2: Block Storage — Sizing and Type
Allocated versus used. A 100 GB volume holding 10 GB bills for 100 GB. Check the gap across the estate.
Shrinking is not an in-place operation. The round trip is: snapshot → create a smaller volume from the snapshot → attach → verify → delete the original. Because it requires downtime and care, it is usually worth doing only where the gap is large.
Volume type migration is the easier win. Newer general-purpose volume types are typically cheaper per GB and provide a better baseline of throughput than the generation before them. For most workloads it is a straightforward change with a performance improvement attached — one of the few optimizations that is unambiguously better on both axes.
Topic 3: Object Storage — Lifecycle Tiering
Object storage grows without anyone deciding it should. Logging buckets, backup buckets and artifact buckets accumulate indefinitely because nothing ever deletes anything.
Lifecycle policies transition objects automatically as they age:
Standard → Infrequent Access → Archive → Deep Archive → Expire
0d 30d 90d 180d 365d
A workable default: not accessed in 30 days → infrequent access tier; not accessed in 90 days → archive. Tune from your own access patterns.
Two things to check before setting an aggressive policy:
- Retrieval cost and latency. Archive tiers are cheap to store and expensive to read, sometimes with retrieval times measured in hours. Data you might need during an incident does not belong in deep archive.
- Minimum storage duration. Cheaper tiers often bill a minimum retention period. Transitioning an object that gets deleted a week later can cost more than leaving it in the standard tier.
Expiry is the lever people avoid because deletion feels irreversible. It is also the only one that stops unbounded growth. Establish the actual retention requirement with whoever owns the data, then encode it.
Topic 4: Network — The Charges Nobody Predicts
Network cost is the layer that produces the most surprise, because the charges do not map to anything you created deliberately.
Managed NAT gateways are the most common surprise. They bill both an hourly rate and a charge per GB processed. A cluster pulling container images and talking to object storage through NAT generates substantial per-GB charges for traffic that never needed to leave the network at all.
The fix is private endpoints. Routing object storage and other provider-service traffic through a private endpoint keeps it internal, bypassing the gateway and its per-GB charge entirely. On a busy cluster this is frequently the single largest network saving available.
Cross-zone data transfer is billed. Traffic between availability zones costs money in a way traffic within one does not, and a chatty microservice architecture spread evenly across three zones generates a great deal of it.
The trap: consolidating zones reduces this cost and reduces your fault tolerance. Never compromise availability topology for cost without explicit business sign-off. This is one of the few optimizations that can turn a cost engagement into an incident review.
Idle public addresses. An allocated address not attached to a running resource bills continuously. Audit and release. Same category as unattached volumes — pure waste, no trade-off.
Load balancers bill on capacity units as well as hours. Several underused load balancers each carrying a fraction of their capacity cost more than one consolidated one.
Inter-region transfer and acceleration services are expensive per GB. For an acceleration service specifically, verify that the latency improvement justifies the cost against a cheaper content delivery approach — the answer is often no.
Topic 5: The Peering Decision
Connecting networks is a cost decision as well as an architectural one:
- Direct peering is cheaper for simple topologies and scales poorly — connections grow quadratically with the number of networks.
- A transit hub adds an hourly charge plus per-GB processing, and collapses that complexity to linear.
For three networks, peering wins. For fifteen, the hub is both cheaper to operate and dramatically simpler. The crossover depends on your traffic volume, and it is worth calculating rather than assuming.
Topic 6: The Sequence
Work these layers in waste-first order, because the early steps carry no risk and fund the conversation for the later ones:
- Delete unattached volumes and release idle addresses. No risk, immediate saving.
- Set snapshot retention. No risk once the retention period is agreed.
- Migrate volume types. Low risk, better performance.
- Add lifecycle policies to buckets that grow unbounded.
- Add private endpoints for provider-service traffic.
- Review cross-zone traffic and peering topology — highest risk, needs architectural agreement.
Try it yourself: List every volume in an unattached state and sort by age. The oldest one is usually measured in years, and it makes the case for a recurring audit better than any policy document.
Common mistake: Setting an aggressive archive lifecycle on a bucket holding logs you might need during an incident. You will save money every month and pay for it once, at the worst possible moment, waiting hours to retrieve the logs that explain an outage.