A node autoscaler provisions nodes to satisfy what pods request. It has no idea whether those requests bear any relationship to what the pods actually use. That gap is where most Kubernetes waste hides, and closing it needs a different class of tool.
Topic 1: Why an Autoscaler Is Not Enough
The clearest framing of the distinction:
- The autoscaler responds to declared demand. A developer declares 4 CPU and 8 GB; it provisions nodes that satisfy exactly that. It is working perfectly and has no visibility into whether the declaration was accurate.
- A cost observability tool compares actual utilisation against what was requested and flags the gap — this workload requested 6 CPU and is using 0.5.
The autoscaler optimises provisioning against a request. The cost tool tells you whether the request itself was right. Genuinely different functions on genuinely different data, which is the answer to “why do I need both”.
There is a second reason: an autoscaler operates in real time within one cluster with no historical layer. Cost tools typically offer a single pane across many clusters with trend data — the view you need to answer “is this getting better or worse?”
Topic 2: What Cluster Cost Tools Recommend
Three recurring recommendation categories, with representative impact from a demonstration cluster spending ~$27,000/month:
| Category | What it finds | Example impact |
|---|---|---|
| Interruptible migration | On-demand workloads that could safely run on Spot | ~$5,000/month |
| Workload right-sizing | Requests materially above actual usage | ~$5,000/month |
| Hibernation | Namespaces or workloads unused at known hours | Proportional to idle time |
Two safety behaviours worth confirming in any tool you evaluate:
- Stateful awareness. A good tool recognises workloads with persistent volumes and will not recommend moving them to interruptible capacity. That is a genuine safety feature, not a nicety.
- Trend before action. Recommendations may appear in real time; do not act on them immediately. Wait for at least a month of trend data. A workload may genuinely need high CPU at a time the current window has not yet seen.
Topic 3: The Landscape
| Tool class | Model | Notes |
|---|---|---|
| SaaS cluster cost platform | Paid, agent reports to vendor | Broadest features; multi-cluster; recommendations plus automation |
| Open-source right-sizing | Free, in-cluster | Request and limit recommendations per workload, driven by vertical-autoscaler metrics |
| Open-source cost visibility | Free, in-cluster | Cost attribution per namespace, workload and pod |
| Observability suite with cost features | Paid, enterprise | Broader scope — traces, logs and cost together; heavier |
| Cloud-native recommender | Free | Covers instances and clusters; does not see inside pods well |
The open-source right-sizing option is a genuinely useful free starting point — one team used it to find services requesting 6 GB of memory and using a fraction of it, which is a very common pattern in JVM workloads where the request was set from the heap size somebody guessed at.
On pricing: SaaS platforms in this space commonly land around $1,000/month or more depending on cluster and workload count. The decision is arithmetic — saving $10,000/month for $1,000 is obviously worth it; saving a few hundred is not.
Topic 4: Where KEDA Fits
Frequently confused with node autoscaling; it operates at a different layer.
KEDA is not a node autoscaler. It creates a ScaledObject which drives a standard horizontal pod autoscaler, which scales pods — and that pod scaling is what then triggers a node autoscaler to add capacity.
event source (queue depth, stream lag, custom metric)
→ KEDA ScaledObject
→ Horizontal Pod Autoscaler
→ more pods
→ node autoscaler provisions nodes
What it adds: the default pod autoscaler scales on CPU and memory. KEDA extends the trigger set to arbitrary event sources — queue depth, stream lag, business metrics. That lets you scale on the signal that actually predicts load rather than on the resource consumption that lags it.
It also supports holding incoming requests briefly while capacity is provisioned, which smooths the gap between a demand spike and nodes becoming available.
The general lesson: platform defaults often need purpose-built extensions for real-world requirements. KEDA and a node autoscaler are complementary layers, not competitors.
Topic 5: The Five-Criterion Selection Framework
A general framework for any infrastructure tool decision, in order of consideration:
1 — Functional fit. Does it actually solve the stated objective? Surface-level capability is easy to confirm; a proof of concept is what reveals differences in response time and decision quality.
2 — Cost. The direct price of the tool.
3 — Non-functional requirements, especially security and compliance. The most consequential and most under-considered category:
- Data residency. If the organisation is legally required to keep data in a region and the SaaS tool only runs servers elsewhere, that alone disqualifies it. Not a negotiation — a disqualification.
- Regulatory frameworks. In regulated sectors, any tool in the environment must itself satisfy the standards the organisation is held to.
The contrast is sharpest exactly here. An in-cluster open-source autoscaler is free and no data leaves your environment. A SaaS cost platform costs money and transmits cluster and workload data to a third party. For a data-sensitive or regulated organisation that is a genuine, potentially disqualifying difference — not a footnote.
4 — Performance impact. Does running the tool degrade what it monitors?
5 — Ease of operation. Does the team have — or can it realistically acquire — the skills to run this well? An excellent tool the team cannot operate is a poor choice.
Work through them in that order. Teams routinely evaluate criterion 1, glance at 2, and discover 3 during a security review after committing.
Topic 6: Recognising Workload Classes
Underpinning every tooling decision is knowing what kind of workload you have:
- Analytics workloads operate after the fact — reporting, recommendation, batch aggregation. Interruption-tolerant, relaxed deadlines, strong candidates for cheap capacity and aggressive right-sizing.
- Real-time workloads process live requests. Lower interruption tolerance, tighter latency requirements, conservative capacity strategy.
Both exist in nearly every business. Classifying a workload correctly is the first step in choosing its node pool, its capacity type and how aggressively to right-size it — and it is a question for the business, not an assumption for the platform team.
Try it yourself: For ten workloads, compute requested versus actual CPU and memory over 30 days and total the gap. That number is what a right-sizing tool would find, and it tells you whether a paid tool is worth evaluating at all.
Common mistake: Buying a cluster cost platform before fixing resource requests. The tool will correctly report enormous waste, you will implement its recommendations, and you will have paid a subscription to discover something a free in-cluster right-sizing tool would have told you. Establish the size of the problem first.