GKE Day-2: Upgrades, Channels and Cost

Upgrades happen whether you plan them or not — release channels, maintenance windows, surge behaviour, and the Autopilot-versus-Standard cost question answered with numbers.

advanced 24 min lesson hands-on task included

A GKE cluster upgrades itself. The only real decisions are how fast, during which hours, and whether your workloads can be drained — and all three are set long before the upgrade happens.


Topic 1: Release Channels

GKE UPGRADES HAPPEN WHETHER YOU PLAN THEM OR NOT — CHOOSE THE CHANNEL Rapid newest, days after release test clusters only Regular the default balance most workloads Stable longest soak time production, risk-averse Extended longer support window slow-moving compliance SURGE UPGRADE — HOW NODES ROLL maxSurge: 1 extra node created first maxUnavailable: 0 never lose capacity Higher surge = faster and more expensive. Zero unavailable is the setting that keeps a rollout invisible. WHAT BLOCKS A NODE UPGRADE · a PDB that allows zero disruptions · a pod with no controller (bare pods never drain) · long terminationGracePeriodSeconds × many pods Same drain mechanics as any Kubernetes upgrade. MAINTENANCE WINDOWS AND EXCLUSIONS ARE THE CONTROL YOU ACTUALLY HAVE gcloud container clusters update … --maintenance-window-start/--end/--recurrence · add an exclusion for your freeze period Autopilot upgrades nodes for you and honours the same windows — you give up node control, not upgrade control.
Pick a channel per cluster and let it manage versions. The panel on the right is the list of things that block a node upgrade, and it is the same drain mechanics as any Kubernetes cluster.
gcloud container clusters update prod --release-channel=stable
ChannelLag behind releaseUse for
RapidDaysTest clusters, early validation
RegularWeeks — the defaultMost workloads
StableMonthsProduction, risk-averse teams
ExtendedLongest support windowSlow-moving compliance environments

Being on a channel is better than pinning a version. A pinned cluster eventually leaves support, and the forced upgrade then jumps several minor versions at once — which is a project rather than a maintenance event.

Run staging one channel ahead of production. Staging on Regular and production on Stable means your workloads meet each version weeks before production does, which is the cheapest possible upgrade testing.

Control plane and nodes upgrade separately. The control plane goes first, automatically; nodes follow according to your node pool settings. Kubernetes supports a version skew of up to two minors between them, so nodes lagging briefly is normal and nodes lagging for months is not.


Topic 2: Maintenance Windows and Exclusions

gcloud container clusters update prod \
  --maintenance-window-start="2026-01-05T02:00:00Z" \
  --maintenance-window-end="2026-01-05T06:00:00Z" \
  --maintenance-window-recurrence="FREQ=WEEKLY;BYDAY=SA,SU"

# A freeze period — Black Friday, an audit, a launch
gcloud container clusters update prod \
  --add-maintenance-exclusion-name=peak-trading \
  --add-maintenance-exclusion-start=2026-11-20T00:00:00Z \
  --add-maintenance-exclusion-end=2026-12-02T00:00:00Z \
  --add-maintenance-exclusion-scope=no_minor_or_node_upgrades

Three exclusion scopes, and the difference matters:

  • no_upgrades — nothing at all, maximum 30 days.
  • no_minor_upgrades — patch versions still apply, maximum 180 days.
  • no_minor_or_node_upgrades — the full freeze for a peak period, maximum 30 days.

An exclusion is a delay, not a cancellation. Everything deferred lands when the window opens, so a long freeze produces a large upgrade immediately afterwards. Plan the catch-up rather than being surprised by it.


Topic 3: Surge Upgrades and What Blocks Them

gcloud container node-pools update apps --cluster=prod \
  --max-surge-upgrade=2 --max-unavailable-upgrade=0
  • maxSurge — extra nodes created before old ones are drained. Higher is faster and briefly more expensive.
  • maxUnavailable — how many nodes may be down at once. Zero is the setting that keeps capacity constant through the rollout.

The sequence per node: cordon, drain (respecting PodDisruptionBudgets and graceful termination), delete, replace.

What stalls a node upgrade, in order of frequency:

CauseSymptom
A PDB allowing zero disruptionsNode stuck SchedulingDisabled, drain never completes
A bare pod with no controllerNothing recreates it, so nothing will evict it
Long terminationGracePeriodSeconds × many podsDrain takes hours and looks stuck
No spare capacity to reschedule ontoPods Pending, drain waits
PVC bound to a zonal disk in the drained zonePod cannot reschedule elsewhere
# Find the PDBs that will block you, before you start
kubectl get pdb -A -o custom-columns=NS:.metadata.namespace,NAME:.metadata.name,ALLOWED:.status.disruptionsAllowed \
  | awk '$3=="0"'

That one command is the pre-flight check worth running before every upgrade, and it is the same finding as the Kubernetes module’s upgrade project.

Blue-green node pool upgrades are the safer alternative for critical pools: GKE creates a whole new pool, shifts workloads, and keeps the old pool available for a fast rollback. Slower and more expensive during the transition, and the rollback is the point.


Topic 4: Autopilot vs Standard, With Numbers

AutopilotStandard
You pay forPod requestsNodes, whether packed or idle
Node managementGoogle’sYours
DaemonSetsLimitedYes
Privileged pods, host accessNoYes
Bin-packingGoogle’s problemYour problem
Right forMost application workloadsAnything needing node-level control

The economics are not obvious in either direction. Autopilot charges a premium per pod but eliminates idle node capacity; Standard is cheaper per unit if — and only if — your nodes are well packed. A Standard cluster at 30% utilisation is almost always more expensive than the same workload on Autopilot.

The honest way to decide is to measure, which is what the hands-on task asks for. Run the same workload both ways for a week and compare the bill. Teams consistently guess wrong.

Cost levers on Standard, in order of effect:

1. Right-size requests. Requests are what the scheduler packs, so
   over-requesting is directly wasted node capacity.
2. Spot node pools for interruptible workloads — 60–91% off, with
   a taint and a 30-second termination notice to handle.
3. Cluster autoscaler with a sane maximum, plus node auto-provisioning.
4. Committed use discounts once the baseline has stopped moving.
5. Delete the idle clusters. There is always at least one.

The Cloud Cost Optimization path covers the general discipline; the GKE-specific point is that requests are the currency — an autoscaler cannot tell you a request was wrong, it can only buy nodes to satisfy it.


Topic 5: The Managed Add-ons Worth Knowing

GKE ships a lot of optional machinery. The ones that change operations:

  • Workload Identity — pods get GCP IAM without any key file. This should be on for every cluster; it is covered in the GKE architecture lesson and it is the single most important security setting.
  • Gateway API controller — the successor to the Ingress controller, and what you need for real weighted traffic splitting (the canary lesson in the CI/CD path assumes it).
  • Backup for GKE — scheduled backup of workload state and PVCs, with restore into another cluster. The thing most teams discover they wanted after an incident.
  • Config Sync / Policy Controller — GitOps and OPA-based policy as managed components, if you would rather not run Argo and Gatekeeper yourself.
  • Managed Prometheus — Prometheus-compatible metrics without running the storage. Integrates with Cloud Monitoring alerting.
  • Cost allocation — per-namespace and per-label cost breakdown in the billing export, which is how you answer “which team is this”.
gcloud container clusters update prod --enable-cost-allocation

Topic 6: An Upgrade Runbook

BEFORE
□ read the release notes for every version you are crossing
□ check deprecated API usage — this is what actually breaks workloads
     kubectl get --raw /metrics | grep apiserver_requested_deprecated_apis
□ list PDBs with disruptionsAllowed = 0 and fix them
□ confirm spare capacity exists for the surge
□ verify a recent backup, and that you have restored one at some point

DURING
□ upgrade the control plane first, watch the API stay responsive
□ upgrade one non-critical node pool, verify workloads reschedule
□ then the rest, with maxUnavailable=0

AFTER
□ every workload Ready, no Pending pods
□ ingress and load balancer health checks green
□ run the smoke tests you would run after a deploy
□ check for pods that came back on a node with a different topology label

The deprecated-API check is the one that catches real breakage. A minor version removal turns a working manifest into a rejected one, and the failure appears at deploy time weeks later rather than during the upgrade.

Try it yourself: create a PDB with minAvailable equal to the replica count and start a node pool upgrade. The drain hangs indefinitely with no error naming the cause — which is exactly what it looks like at 2am, and why the pre-flight awk one-liner above is worth keeping.

Common mistake: setting a maintenance exclusion for a peak period and forgetting to remove it. Upgrades stop for the maximum the scope allows, the cluster silently drifts out of the channel’s supported window, and the eventual forced upgrade crosses several versions at once — with all the deprecated-API breakage arriving together. Set an end date, and put its removal in the calendar.