Scheduling, Spot & Commitment Discounts

Three levers that need no code change: shutting non-production down out of hours, running interruptible capacity safely in production, and buying commitment discounts without over-committing.

advanced 22 min lesson hands-on task included

Right-sizing changes what you run. These three levers change when you run it, what kind of capacity you run it on, and how you pay for it — none of them require touching the application.


Topic 1: Non-Production Scheduling

Dev, staging and QA typically run 24×7 while the team that uses them works business hours. That is roughly two-thirds of the bill for those environments buying nothing.

Stop at 22:00, start at 08:00 — a 14-hour running window, about a 42% reduction on non-production compute. Tighter windows are possible; this one survives contact with people who occasionally work late.

The mechanism is tag-filtered and the tag filter is the safety mechanism:

import boto3

def lambda_handler(event, context):
    ec2 = boto3.client('ec2')
    response = ec2.describe_instances(
        Filters=[
            {'Name': 'tag:Environment',      'Values': ['dev', 'staging']},
            {'Name': 'instance-state-name',  'Values': ['running']},
        ]
    )
    ids = [i['InstanceId']
           for r in response['Reservations']
           for i in r['Instances']]
    if ids:
        ec2.stop_instances(InstanceIds=ids)
    return {'stopped': ids}

Production never matches the filter, which is why it is never at risk. That property is worth stating in the change record.

Where to host it: a scheduled function is fully managed but adds its own small cost. A cron job on existing CI infrastructure adds none, and is the better choice when you already run CI. One estate deliberately moved scheduler functions onto their build server for exactly this reason.

Set the cron in UTC and convert deliberately — an off-by-one on the timezone stops the environment during the working day, which is a memorable way to lose trust in cost automation.

Honest scoping: this lever pays proportionally to how concentrated your team’s hours are. One organisation with a round-the-clock dev team recovered only ~$900/month from it. Check working patterns before promising a number.


Topic 2: Interruptible (Spot) Capacity

Spare capacity at 70–90% off, reclaimable at roughly two minutes’ notice.

The myth: “Spot is only for dev and test.” The reality: large consumer platforms run substantial production workloads on it. Spot is safe in production if the architecture handles interruption. Bad architecture plus Spot is catastrophic; good architecture plus Spot is a large saving at acceptable risk.

Suits SpotWhy
Dev, QA, stagingNo special architecture needed
Stateless servicesSafe with ≥2 replicas, autoscaling and a disruption budget
Queue consumersInterruption returns the message to the queue
Batch and ETLJob retry handles it
CI build agentsBuild reruns; no state
Distributed data processingNative interruption support in most engines
Risky on SpotWhy
Stateful services — databases, brokers with local storageState is lost; recovery is complex
Single-replica critical servicesInterruption is 100% downtime
Anything without retryInterruption means lost work

The production pattern — a mixed pool:

Cluster
├── Pool A — ~70% interruptible, DIVERSIFIED across instance types
│     t3a.medium, t3.medium, m5a.large, m5.large …
└── Pool B — ~30% on-demand, stable baseline
The production mix — split at the node layer, not the application layer ~70% interruptible DIVERSIFIED — this is the part people skip t3a.medium t3.medium m5a.large m5.large batch · queue consumers · CI agents · stateless services one instance type only = no capacity when that type is reclaimed ~30% on-demand stable baseline always-available capacity stateful services single-replica critical paths Required safeguards before any of this is production-safe ≥2 replicas · PodDisruptionBudget · horizontal autoscaling · SIGTERM handling · multi-zone spread You do not assign services to pools. The scheduler places pods on any ready node; the mix guarantees fallback capacity exists.
The mix is a property of the node layer. Diversification across instance types is what stops a single reclamation event emptying the pool, and the safeguards along the bottom are what make any of it production-safe.

Diversification is the part people skip and it is the part that matters. If the interruptible pool only requests one instance type and the provider reclaims that type in your zone, you have no capacity at all. Spreading across 3–5 compatible types makes simultaneous reclamation dramatically less likely.

The split is on nodes, not on applications

A point that causes real confusion. You do not assign specific services to interruptible capacity and others to stable capacity. The scheduler places pods on any ready node.

What the 70/30 mix guarantees is that stable fallback capacity always exists. When interruptible nodes are reclaimed, their pods reschedule onto the on-demand nodes automatically. The ratio is a property of the cluster, not an assignment of workloads.

Taints and tolerations are how you make exceptions to that — pinning stateful or single-replica services to stable capacity — not how you implement the ratio itself.

Required safeguards before any of this is production-safe

SafeguardWhat it does
≥2 replicas per workloadSingle replica plus interruptible capacity is guaranteed downtime on reclamation
PodDisruptionBudgetLimits how many pods can be disrupted at once, so a drain cannot take every replica
Horizontal pod autoscalingRestores capacity when a node disappears
Graceful shutdownThe application catches the termination signal and finishes in-flight requests inside the ~2-minute window
Instance type diversificationAvoids simultaneous reclamation of the whole pool
Multi-zone spreadA zone-level capacity shortage does not take down every pod

The graceful-shutdown row is the one most often missing. Without signal handling, the two-minute warning buys you nothing — the process dies mid-request and clients see errors that look like an application bug rather than a capacity event.

Choosing by workload category:

A useful lens: analytics workloads operate after the fact — reports, recommendations, batch aggregation. They tolerate interruption and relaxed deadlines, and are excellent Spot candidates. One analytics estate ran a $40–50k/month processing workload almost entirely on interruptible capacity, where on-demand would have added roughly 30% for no benefit.

Real-time workloads process live requests and need lower interruption tolerance. Both categories exist in nearly every business, and classifying a workload is the first step in deciding its capacity strategy.

A caution on deadlines: the acceptable delay for a batch job is set by the business consumer of its output, not by the infrastructure team’s assumption. A finance report might tolerate a week’s delay most of the year and almost none at fiscal year end. The same job’s criticality is not fixed — it varies by calendar. Ask; do not assume.


Topic 3: Commitment Discounts

Commit to a level of spend for one or three years in exchange for a discount.

Two shapes:

ShapeCoversFlexibilityTypical discount
Flexible compute commitmentInstances across any family, size, region and OS — plus serverless functions and managed container computeHighest~17–66%
Instance-family commitmentOne instance family in one regionLower~20–72%

The coverage difference is the one that costs people money. The flexible plan covers serverless functions and managed container compute; the family-specific plan does not. If your stack includes either — and most do — the family plan silently leaves that spend at full price.

The safe default is the flexible plan. Choose the family-specific one only when you are genuinely stable on one family in one region and have modelled the difference. The most common purchasing mistake is picking the wrong type, not the wrong amount.

TermPaymentApproximate discount
1 yearNo upfront~20–30%
1 yearPartial upfront~25–35%
1 yearAll upfront~30–40%
3 yearsAll upfront~40–60%

The rules that matter:

Commitments apply to on-demand capacity, not interruptible capacity. Buying a commitment expecting it to cover Spot usage buys nothing. Decide your Spot strategy first, then commit against what remains.

This makes the two levers complementary rather than alternatives, and it is the cleanest way to think about the whole topic:

  • Interruptible capacity handles your volatile, tolerant workloads cheaply.
  • Commitments handle your stable, always-on workloads cheaply.

Every workload belongs to one of those two categories. Classify first, then apply the matching lever — and never commit against spend that Spot is about to remove.

Commit to the baseline, never the peak. Unused commitment is money spent for nothing. Find the floor your spend never drops below and commit to a portion of that. Under-committing leaves some savings unclaimed; over-committing is a loss you carry for up to three years.

There is no exit. Commitments cannot be cancelled and generally have no secondary market. This is a one-to-three-year decision made on data you gathered over ninety days — treat it with corresponding seriousness.

Do not buy on the recommendation engine alone. The provider’s calculator is a useful input, not a decision. Understand the arithmetic behind the number: what your steady-state floor is, how confident you are that it holds, and what happens to the commitment if you migrate architecture or move a workload to a different platform next year.

That last point interacts directly with the previous lesson. Migrating to a cheaper CPU architecture reduces your on-demand spend — so a commitment bought before that migration may end up over-committed afterwards. Sequence the work: migrate first, re-baseline, then commit.


Try it yourself: Chart your daily compute spend for 90 days and draw a line at the minimum. That floor is your defensible commitment level. Almost everyone’s first instinct is to commit well above it.

Common mistake: Buying a three-year all-upfront commitment early in a cost engagement because it shows the largest headline saving. Every subsequent optimization you perform reduces the spend that commitment was sized against — and you have locked in the pre-optimization baseline for three years.