Cloud Cost Optimization
Cloud spend as an engineering discipline β the layer model, right-sizing without causing outages, Kubernetes node economics, and the observability bill nobody checks.
Stage 1 β Foundations
3 lessonsWhy cloud cost is an engineering discipline rather than a finance report, the layered framework that stops you missing whole categories of spend, and the signal that tells you where to start.
The three foundations that must exist before any optimization: a tag schema that makes cost attributable, a full resource inventory that doubles as an audit trail, and a utilization baseline long enough to be trusted.
The four provider-native tools to enable before any optimization work, what each is actually good for, how they differ across clouds, and using cost explorers to prove a change worked.
Stage 2 β Compute
3 lessonsThe three-pillar rule that stops cost work causing outages, sizing against P99 rather than average, the buffer that absorbs spikes, and the four situations where right-sizing should be deferred.
Moving between CPU vendors and instruction sets for 15β40% savings: why one hop is nearly free and the other needs an assessment, the inverted complexity matrix across instance categories, and the migration and rollback procedure.
Three levers that need no code change: shutting non-production down out of hours, running interruptible capacity safely in production, and buying commitment discounts without over-committing.
Stage 3 β Kubernetes Cost
3 lessonsDeriving min, desired and max capacity from business numbers rather than guesswork, the 20% buffer rule, and the three structural problems that make classic node autoscaling overprovision by design.
Provisioning the cheapest node that fits the pod instead of scaling a fixed pool: the six-step setup, EC2NodeClass and NodePool, how consolidation recovers off-peak cost, and node pool design patterns that hold up in production.
Why an autoscaler cannot tell you a request was wrong, what utilisation-observability tools add, where KEDA fits, and a five-criterion framework for choosing infrastructure tooling that treats compliance as a first-class factor.
Stage 4 β Beyond Compute
2 lessonsThe two layers people check last: unattached volumes and snapshot sprawl, volume type migration, object lifecycle tiering, and the network charges β gateways, cross-zone transfer, idle addresses β that never appear as a line you recognise.
The layer that is 44% of the bill on some estates: separating retention from volume, distinguishing audit logs you cannot touch from product logs you can, and cutting spend without losing the data you debug with.
Stage 5 β Capstone Project
1 lessonRun a complete engagement end to end β inventory and baseline, layer-by-layer analysis, a sequenced plan ordered by risk, implementation with verification, and the reporting that makes savings defensible.
πΊοΈ Beginner β Expert Roadmap
5 stages with prerequisites and a concrete mastery check at each.
π― What You'll Learn
- β’ Read a bill as a diagnostic signal, and let the billing data choose which layer to attack first.
- β’ Build the tag schema, inventory and 90-day baseline that every later decision depends on.
- β’ Right-size against P99 with a buffer, and name the four situations where the correct action is to wait.
- β’ Migrate CPU architecture for 15β40% savings, and explain why the difficulty inverts between instance categories.
- β’ Schedule non-production shutdowns safely, using the tag filter as the mechanism that protects production.
- β’ Decide where interruptible capacity belongs in production, and diversify a Spot pool so it survives reclamation.
- β’ Size a commitment against your baseline floor β and know why buying one early is the mistake you cannot undo.
- β’ Derive cluster max capacity from business numbers rather than a round figure someone typed.
- β’ Run Karpenter with node pools that give it room to optimise, and set the ceiling that bounds a runaway.
- β’ Explain why an autoscaler cannot tell you a resource request was wrong, and what closes that gap.
- β’ Apply a five-criterion tool-selection framework where data residency can disqualify a tool outright.
- β’ Cut storage, network and log spend β the three layers where waste accumulates without anyone deciding.
- β’ Run a full engagement in risk-ascending order and report savings with per-day, per-resource evidence.
π‘οΈ Best Practices in Production
The short version of this path. Every lesson also ends with the specific mistake it exists to prevent.
- β Start from the bill: let billing data choose which layer to attack, not intuition.
- β Fix tagging and build a 90-day baseline before optimising anything.
- β Right-size against P99 with a buffer, and change one variable at a time.
- β Work in risk-ascending order: waste first, then low-risk changes, then architecture.
- β Buy commitments last, once the baseline floor has stopped moving.
- β Diversify Spot pools across instance types and zones, and handle the interruption notice.
- β Attack storage, network and log retention β the three layers where waste accumulates unnoticed.
- β Report savings with per-day, per-resource evidence and a tested rollback for every change.
- β Buying a Savings Plan or RI before right-sizing. You have committed to your own waste.
- β Right-sizing on average utilisation, which turns a cost win into a latency incident.
- β Turning off 'idle' resources without checking who owns them or what depends on them.
- β Treating an autoscaler as evidence that requests are correct β it cannot tell you that.
- β Cutting log retention or sampling without asking what an incident investigation needs.
- β Reporting a projected saving that never appears in a bill because nobody verified it.
πΌ Interview Readiness
Once you reach the end of this path, test your engineering knowledge against real questions asked by top technical teams:
A client has a $98K/month AWS bill. Walk me through how you'd approach reducing it.
Cloud Cost OptimizationA client wants to know why their cloud bill increased significantly compared to last month. How would you approach answering this credibly?
Cloud Cost OptimizationHow would you explain to a non-technical stakeholder (e.g., a compliance officer) the difference between "we stopped storing these logs in our operational bucket" and "we deleted this data," in a way that would satisfy an audit conversation?
Cloud Cost OptimizationWhy is resource tagging described as foundational to cost optimization work?
Cloud Cost OptimizationWhat tags should every AWS resource have, and why?
Cloud Cost OptimizationDesign a systematic inventory-building process for a GCP account with no existing documentation, using the gcloud CLI. What categories would you capture, and why does the organizational structure matter?
Cloud Cost OptimizationHow do you implement FinOps for a multi-account AWS organization with 20 accounts?
Cloud Cost OptimizationWhat is AWS Cost Explorer and why is it the first tool you enable for cost optimization?
Cloud Cost OptimizationWhat is AWS Trusted Advisor, and what categories of recommendations does it provide?
Cloud Cost OptimizationWhy is GCP's Active Assist described as a meaningful differentiator versus AWS Trusted Advisor?
Cloud Cost OptimizationWhat is a GCP SKU, and why would you filter billing data by it?
Cloud Cost OptimizationExplain the PRC framework for right-sizing decisions.
Cloud Cost OptimizationExplain why an engineer might deliberately avoid right-sizing instances during a cloud-to-cloud migration, even though both activities individually seem like sound cost-optimization practice.
KubernetesHow would you decide which specific EC2 instances in a large fleet are good candidates for Savings Plan coverage versus Spot versus on-demand, as part of a comprehensive cost-optimization plan?
KubernetesWhat is MaxSessions / right-sizing, and how do you determine the correct instance size?
Cloud Cost OptimizationWhat's the difference between Intel, AMD, and ARM/Graviton instance families on AWS, from a cost perspective?
Cloud Cost OptimizationWhy does the recommended migration path go Intel β AMD β ARM instead of directly Intel β ARM?
Cloud Cost OptimizationExplain the difference in migration complexity between a standalone EC2 instance, an instance in an Auto Scaling Group, and an EKS node group instance β and why does the difficulty ordering change depending on migration direction?
Cloud Cost OptimizationWalk through an ASG IntelβAMD migration with zero downtime.
Cloud Cost OptimizationWhy does EKS node group migration require ~2β3 minutes of downtime, while ASG migration has zero downtime?
Cloud Cost OptimizationWhy can you NOT specify an IAM instance profile in the Launch Template when creating EKS node groups via CLI?
Cloud Cost OptimizationWhen is it appropriate to stop EC2 instances to save cost, and when is it not?
Cloud Cost OptimizationA team wants to use spot instances for a production Kubernetes workload but is worried about interruption risk. What architectural safeguards would you recommend?
KubernetesWhat are the mandatory Kubernetes safeguards before running Spot instances in production?
Cloud Cost OptimizationExplain the tradeoffs between Compute Savings Plans, EC2 Savings Plans, and Reserved Instances.
Cloud Cost OptimizationWhat are the risks of buying a Savings Plan without doing the mathematics first?
Cloud Cost OptimizationHow would you decide whether a given batch workload is a good candidate for Spot instances?
Kubernetes AutoscalingCompare capacity planning approaches: traditional (static max) vs. Karpenter-managed. When is each appropriate?
Kubernetes AutoscalingWhat is a cluster autoscaler, and why does a Kubernetes cluster need one?
Cloud Cost OptimizationWhat is the difference between Karpenter and the Kubernetes Cluster Autoscaler?
Kubernetes AutoscalingA namespace has no ResourceQuota/LimitRange, and a BestEffort batch job is co-located with a Burstable production service on the same node group. What's the risk, and how would you redesign this?
KubernetesWalk through the 6 steps to install Karpenter on an EKS cluster.
Kubernetes AutoscalingHow does Karpenter know which subnets and security groups to use when provisioning a new node?
Kubernetes AutoscalingWhat does Karpenter's consolidation feature do and why does it save cost?
Kubernetes AutoscalingKarpenter provisioning is taking longer than expected during a traffic spike. What would you investigate?
Kubernetes AutoscalingDesign a node pool strategy for an EKS cluster that needs to support both a real-time customer-facing calculation service and a large nightly Spark analytics job, using the principles discussed in this guide.
Kubernetes AutoscalingWhat are the main risks or downsides of over-customizing Karpenter's NodePool configuration for every workload type upfront?
Kubernetes AutoscalingWhat is CastAI and how is it different from Karpenter?
Kubernetes AutoscalingA team says "Karpenter already optimizes our costs, so we don't need a separate cost-monitoring tool." How would you respond?
Kubernetes AutoscalingCompare CastAI, Karpenter, and a standard Kubernetes cluster-autoscaler. When would you recommend each?
Cloud Cost OptimizationA prospective client is based in the European Union and has strict data-residency requirements. You're evaluating CastAI (US-only SaaS) versus Karpenter (open-source, in-cluster) for their EKS autoscaling needs. Walk through your decision process.
Kubernetes AutoscalingHow would you use the "functional vs. non-functional requirement" distinction to structure a broader infrastructure tool evaluation (not just autoscaling), and why does treating security/compliance as a distinct, dedicated evaluation category matter?
Kubernetes AutoscalingWhat is an S3 (or GCS) lifecycle policy, and what problem does it solve?
Cloud Cost OptimizationExplain the trade-off between using a managed cloud service versus self-hosting the equivalent open-source tool (e.g., a managed message queue vs. self-hosted Kafka).
Cloud Cost OptimizationYou're told a GCP account's logging costs are unexpectedly high. Walk through how you'd investigate and address it.
Cloud Cost OptimizationWhat's the difference between GCP's _Default and _Required log buckets?
Cloud Cost OptimizationWhat's the difference between reducing log retention and excluding log categories, as two separate cost-optimization techniques?
Cloud Cost OptimizationA client refuses to let you make application-level logging changes, but wants logging costs reduced. What are your levers, and what are the trade-offs of each?
Cloud Cost OptimizationWhy might a compliance/audit log bucket be configured to retain data for 400 days even when the regulatory minimum is only 365 days?
Cloud Cost OptimizationWhat's a quick, low-risk way to reduce Prometheus-related cost without removing any metrics?
Cloud Cost OptimizationDesign a sequencing strategy for a multi-phase cloud cost optimization engagement, using the principle demonstrated in this guide (start with the lowest-risk changes).
Cloud Cost OptimizationHow would you structure a cost optimization engagement for a client with no existing infrastructure documentation?
Cloud Cost OptimizationDesign a systematic process for auditing and reducing an organization's Cloud Monitoring/observability costs, using the approach demonstrated in this guide as a starting point.
Cloud Cost OptimizationHow would you reconcile a scenario where the aggregate billing figure for a service (e.g., "30,000 GB over 90 days") doesn't cleanly match a per-day calculation presented separately (e.g., "30 GB/day Γ 30 days")? What would you do before presenting a cost-savings number to a client?
Cloud Cost OptimizationWhat's the risk profile of making infrastructure changes to a system with no clear ownership or point of contact, and how do you mitigate it?
Cloud Cost Optimization