Install works. Upgrades are where Helm meets the state that already exists — and where most of its surprising behaviour lives. Almost all of it follows from one mechanism.
Topic 1: The Three-Way Merge
On upgrade, Helm computes a patch from three inputs:
- The old manifest — what revision N applied, from the release Secret.
- The live object — what is in the cluster right now.
- The new manifest — what your chart renders today.
The patch contains only what actually needs to change, and it is applied to the live object.
This is why a kubectl edit disappears. Suppose someone scaled a Deployment by hand:
| Value | |
|---|---|
| Old manifest (rev N) | replicas: 3 |
| Live object | replicas: 8 ← edited by hand |
| New manifest | replicas: 3 |
Helm sees that the chart still says 3, that the live value differs, and reconciles it back to 3. The edit vanishes at the next upgrade, hours or days later, with no obvious cause. In Helm 2 — which compared only old and new — the edit survived, which is why long-time users find this behaviour surprising.
The practical rule: any change that should persist belongs in the chart or in the values. For scaling specifically, either use an HPA (in which case the chart should not set replicas at all, or it will fight the autoscaler) or change the value and upgrade.
The corollary for a chart author: do not template a field that a controller owns. A chart that sets replicas while an HPA manages the same Deployment produces a fight that shows up as replica counts oscillating on every deploy.
Topic 2: Immutable Fields
Some fields cannot be changed after creation. A patch containing one is rejected outright, and no Helm flag makes it work:
| Object | Immutable |
|---|---|
| Deployment, StatefulSet, DaemonSet | .spec.selector |
| Job | .spec.template (almost all of it), .spec.selector |
| PersistentVolumeClaim | storageClassName, accessModes, size decreases |
| Service | .spec.clusterIP, and type in some transitions |
| StatefulSet | Most of .spec except replicas, template, updateStrategy |
Error: UPGRADE FAILED: cannot patch "checkout" with kind Deployment:
Deployment.apps "checkout" is invalid: spec.selector: Invalid value: ...
field is immutable
Selectors change more often than people expect — adding a label to selectorLabels, renaming the chart, or changing a fullnameOverride. That is exactly why the anatomy lesson insisted that selectorLabels contain nothing that varies with version.
The recovery options, in order of preference:
# 1. Revert the change. Usually the right answer for an accidental selector change.
# 2. Deliberate replacement, in a maintenance window, accepting downtime:
kubectl delete deployment checkout -n prod --cascade=orphan # keep pods serving
helm upgrade checkout . -n prod # recreate
kubectl delete pods -l app.kubernetes.io/instance=checkout,<old-label> -n prod
# 3. For a PVC size decrease or storage class change: create a new PVC,
# migrate the data, and switch — there is no in-place path.
--cascade=orphan is the detail that makes option 2 survivable: it deletes the Deployment object while leaving the pods running, so traffic continues while the new Deployment adopts them.
helm upgrade --force is not the answer. It deletes and recreates resources, which for a Deployment means every pod stops before new ones start. It has narrow legitimate uses and does not belong in a pipeline.
Topic 3: The Five Failures You Will Actually Meet
another operation (install/upgrade/rollback) is in progress
A previous run was interrupted — CI timed out, a runner was evicted, someone pressed Ctrl-C. The release is stuck in pending-upgrade and Helm refuses to start another operation.
helm history checkout -n prod # find the pending revision
# If the previous revision is healthy, rolling back clears the lock:
helm rollback checkout <last-good> -n prod
# If that refuses, remove the pending revision's record:
kubectl get secret -n prod -l owner=helm,name=checkout --sort-by=.metadata.creationTimestamp
kubectl delete secret sh.helm.release.v1.checkout.v7 -n prod
Deleting the Secret removes Helm’s record of that attempt — it does not change anything running. Check the live state first, because after this Helm’s idea of the release comes from the previous revision.
timed out waiting for the condition
--wait expired. Helm is reporting, not causing. Go to the pods:
kubectl describe pod -n prod -l app.kubernetes.io/instance=checkout | sed -n '/Events:/,$p'
kubectl logs -n prod -l app.kubernetes.io/instance=checkout --tail=50 --previous
Usually: image pull failure, failing readiness probe, insufficient quota, or a PVC that will not bind.
A failed hook — read the hook pod’s logs, which exist if you left hook-failed out of the delete policy.
”Nothing changed but a revision appeared” — Helm records a revision for every upgrade whether or not the manifests differ. Harmless, and it is why helm diff upgrade before running is worth the habit.
Topic 4: CRDs Are Outside the Model
Files in crds/ are installed once, before everything else, and then never touched: not upgraded, not deleted, not rolled back.
mychart/
└── crds/
└── widgets.acme.example.yaml
The reasoning is sound — a CRD is cluster-scoped and shared, and an automatic upgrade could break other users of the same type, while a delete would remove every custom resource of that kind. The consequence is that CRD updates are your job:
kubectl apply --server-side -f crds/ # before helm upgrade
helm upgrade checkout . -n prod
Some charts instead ship CRDs as normal templates guarded by a value (crds.install: true), which makes them upgradeable and makes uninstall dangerous. Neither approach is wrong; know which one a chart uses before you upgrade it, because the failure mode of guessing is either “the new field is unknown” or “all your custom resources were deleted”.
Topic 5: Making Upgrades Boring
A production upgrade sequence worth standardising:
# 1. What will change?
helm diff upgrade checkout . -f values/prod.yaml -n prod
# 2. Will the API server accept it?
helm upgrade checkout . -f values/prod.yaml -n prod --dry-run=server > /dev/null
# 3. Do it, with a bounded blast radius and an audit trail
helm upgrade --install checkout . -f values/prod.yaml -n prod \
--atomic --timeout 8m --history-max 10 \
--description "deploy ${GIT_SHA} via ${CI_JOB_URL}"
# 4. Verify beyond "the command exited 0"
helm test checkout -n prod
Chart-side properties that prevent failures rather than surviving them:
- Never template a field a controller owns —
replicaswhen an HPA is enabled,clusterIP, anything a mutating webhook injects. - Keep
selectorLabelsinvariant across versions. before-hook-creationon every hook Job.- Roll pods on config change with the
checksum/configannotation, rather than expecting a ConfigMap edit to restart anything. - Set
PodDisruptionBudgetdeliberately —minAvailableequal to the replica count blocks node drains and makes cluster upgrades fail. This interacts with the Kubernetes upgrade path, covered in that module’s drain lesson. - Make probes reflect readiness, not liveness of a dependency. A readiness probe that checks the database turns a database blip into every pod being marked unready at once.
Topic 6: What —wait Actually Waits For
Worth being precise, because “it said it was ready” is a claim people make about it:
--wait waits for Deployments, StatefulSets and DaemonSets to reach their expected ready replica count, PVCs to be bound, Services to have endpoints (except ExternalName), and Jobs to complete. It does not wait for custom resources unless the chart uses --wait-for-jobs semantics or the CRD reports readiness in a way Helm understands.
Which means a chart that creates a Certificate, a Kafka or another operator-managed resource can report success while the operator has not finished. If readiness of a custom resource matters, wait for it explicitly after the upgrade:
kubectl wait --for=condition=Ready certificate/checkout-tls -n prod --timeout=5m
--wait-for-jobs extends the wait to Jobs that are part of the release (not hooks, which are always waited for). Useful when the release includes a one-shot Job whose completion matters.
Try it yourself: scale a Deployment with kubectl scale, then run helm upgrade with no chart changes at all. The replica count snaps back to the chart’s value. That single experiment explains a whole category of “the cluster keeps undoing my change” reports — and it is why the chart, not the cluster, is where changes belong.
Common mistake: responding to another operation is in progress by deleting the entire release and reinstalling. That deletes running workloads for what is a bookkeeping problem — the release Secret says pending, the cluster is fine. Roll back to the last good revision, or delete the single pending revision Secret. Both preserve the workload; reinstalling is an outage you chose.