Install, Upgrade and Rollback

Revisions as an append-only log, the four flags that make an upgrade survivable, and the honest limits of what a rollback can undo.

beginner 20 min lesson hands-on task included

Install is the easy half. Everything interesting about Helm happens on the second deploy, when there is existing state to reconcile and a decision to make about what happens if the new version does not come up.


Topic 1: Revisions Are Append-Only

EVERY OPERATION APPENDS A REVISION — INCLUDING A ROLLBACK 1 deployed helm install app 1.4.0 1 Secret 2 superseded helm upgrade app 1.5.0 1 Secret 3 failed helm upgrade app 1.6.0 — bad probe 1 Secret 4 deployed helm rollback 2 a COPY of revision 2 1 Secret …5, 6, 7 history-max trims old ones THE FLAGS THAT MAKE UPGRADES SURVIVABLE --atomic roll back automatically on failure --wait --timeout wait for readiness, then decide --install upgrade-or-install, idempotent for CI WHAT ROLLBACK CANNOT UNDO · a database migration a hook already ran · data written by the new version · a deleted PVC, or a CRD change READ THE HISTORY BEFORE YOU ROLL BACK $ helm history checkout -n prod REVISION UPDATED STATUS CHART APP VERSION DESCRIPTION 3 Mon Aug 18 09:12 failed checkout-0.4.1 1.6.0 Upgrade "checkout" failed: timed out 2 Mon Aug 18 08:40 superseded checkout-0.4.0 1.5.0 Upgrade complete ← roll back to this
A rollback does not delete revision 3 — it creates revision 4 containing revision 2's content. History is append-only, which is why `helm history` is the first command to run when something is wrong.
helm install checkout ./chart -n prod            # revision 1
helm upgrade checkout ./chart -n prod            # revision 2
helm upgrade checkout ./chart -n prod            # revision 3 (fails)
helm rollback checkout 2 -n prod                 # revision 4 == revision 2's content
helm history checkout -n prod
REVISION  UPDATED            STATUS      CHART           APP VERSION  DESCRIPTION
1         Mon Aug 18 08:10   superseded  checkout-0.3.9  1.4.0        Install complete
2         Mon Aug 18 08:40   superseded  checkout-0.4.0  1.5.0        Upgrade complete
3         Mon Aug 18 09:12   failed      checkout-0.4.1  1.6.0        Upgrade "checkout" failed: timed out
4         Mon Aug 18 09:20   deployed    checkout-0.4.0  1.5.0        Rollback to 2

The status column is the useful part:

StatusMeaning
deployedThe current, live revision
supersededA former revision, replaced by a later one
failedThe operation did not complete
pending-install / pending-upgrade / pending-rollbackAn operation is running — or crashed partway
uninstallingDeletion in progress

A stuck pending-* status is the single most common Helm operational problem. It means a previous command was interrupted — CI timed out, a laptop closed, a pod running Helm was evicted — and Helm now refuses to start another operation with another operation (install/upgrade/rollback) is in progress. The recovery is in the upgrade-failures lesson; recognising the cause here is what makes it a two-minute fix.

Bound the history, or every upgrade grows etcd:

helm upgrade --install checkout ./chart -n prod --history-max 10

Topic 2: The Flags That Matter

helm upgrade --install checkout ./chart -n prod \
  -f values-prod.yaml \
  --atomic \
  --timeout 5m \
  --history-max 10

--install — install if absent, upgrade if present. Makes the command idempotent, which is what a pipeline needs.

--wait — do not return until all resources report ready: Deployments to their expected replica count, PVCs bound, Services with endpoints, Jobs completed. Without it, helm upgrade returns as soon as the API server accepts the objects, and a green pipeline means nothing about whether the release works.

--timeout (default 5m) — how long --wait waits. Set it above your slowest legitimate rollout; a chart with a database migration hook and a slow image pull needs more than five minutes, and a timeout that fires on a healthy deploy is worse than no timeout.

--atomic — implies --wait, and rolls back automatically if the upgrade fails or times out. This is the flag that turns a failed deploy into a non-event. The cost: the rollback takes time too, so a failing deploy occupies the pipeline for up to twice the timeout.

--cleanup-on-fail — delete resources created during a failed upgrade instead of leaving them orphaned. Worth adding whenever a chart creates new objects between versions.

--dry-run=server — render, send to the API server for validation, apply nothing. Catches schema errors, admission webhook rejections and RBAC problems that client-side rendering cannot.

--force — deletes and recreates resources rather than patching. It is not a stronger --wait; it causes downtime and can destroy things. There is a narrow legitimate use (an immutable field change you have decided to accept downtime for) and it should never be a reflex in a pipeline.

A defensible production command:

helm upgrade --install checkout ./chart \
  --namespace prod --create-namespace \
  -f values/base.yaml -f values/prod.yaml \
  --set image.tag="${GIT_SHA}" \
  --atomic --timeout 8m --history-max 10 \
  --description "deploy ${GIT_SHA} by ${CI_JOB_URL}"

--description is underused: it lands in the helm history DESCRIPTION column, which turns the release log into something an on-call engineer can correlate with a pipeline run.


Topic 3: What Rollback Actually Does

helm rollback re-applies a previous revision’s stored manifest. That is a precise and limited promise, and knowing its limits prevents a bad assumption during an incident.

helm rollback checkout          # to the previous revision
helm rollback checkout 2        # to a specific one
helm rollback checkout 2 --wait --timeout 5m

It reverts: the Kubernetes objects the release manages — image tags, replica counts, environment variables, resource limits, config.

It does not revert:

  • Anything a hook did. A migration Job that ran during the failed upgrade has already changed your database. Helm has no undo for it.
  • Data written by the new version. Rows, files, queue messages, cache entries in a new format.
  • Deleted PVCs, and it will not shrink a volume that was expanded.
  • CRDs, which Helm never upgrades or deletes at all.
  • Anything outside the release — a DNS record, an external load balancer created by a controller, a row in another system.

The consequence for schema migrations is the most important operational rule in this module: migrations must be backward compatible with the previous application version, or you do not have a rollback. Add a column, deploy code that writes both, remove the old column two releases later. If the migration drops a column the previous version reads, rolling back the application produces a broken application, faster.


Topic 4: Uninstall, and What Survives It

helm uninstall checkout -n prod
helm uninstall checkout -n prod --keep-history   # release record stays, resources go
helm list -n prod --uninstalled

What Helm does not remove:

  • PersistentVolumeClaims created by a StatefulSet’s volumeClaimTemplates. Kubernetes owns those, deliberately, so your data survives. They must be deleted by hand — and they keep costing money until somebody does.
  • CRDs from crds/, and therefore every custom resource of those kinds.
  • Anything with helm.sh/resource-policy: keep, which is a per-resource opt-out from deletion.
  • Resources created by hooks, unless a delete policy removed them.
metadata:
  annotations:
    helm.sh/resource-policy: keep

That annotation is the right tool for a PVC or a Secret that must outlive the release, and a trap when applied carelessly — a kept resource blocks a later reinstall with “already exists”.

After any uninstall, check for orphans:

kubectl get pvc,secret,cm -n prod -l app.kubernetes.io/instance=checkout

Topic 5: Inspecting a Live Release

When something is wrong, these four commands answer most questions before you open the chart:

helm status checkout -n prod            # current revision, resources, NOTES
helm get values checkout -n prod        # what values were supplied
helm get values checkout -n prod -a     # ...merged with chart defaults (the real config)
helm get manifest checkout -n prod      # exactly what was applied
helm get hooks checkout -n prod         # hook manifests attached to this release

helm get manifest is the ground truth. It is what Helm actually sent to the API server for the current revision — not what the chart in your working directory renders today. The difference between the two is drift, and comparing them is the fastest way to answer “is production running what is in main?”:

diff <(helm get manifest checkout -n prod) \
     <(helm template checkout ./chart -f values/prod.yaml --namespace prod)

helm get values -a deserves its own mention: without -a you see only the overrides, which hides the defaults that are doing the real work. During an incident, -a is the flag you want.


Topic 6: Deploying the Same Chart Repeatedly

A chart is a package; a release is an installation. Several releases of one chart in one namespace is normal and is a good test of chart quality:

helm install checkout ./chart -n prod
helm install checkout-canary ./chart -n prod -f values/canary.yaml

For this to work, every resource name must derive from .Release.Name, every selector must include app.kubernetes.io/instance, and nothing may use a cluster-scoped fixed name (a ClusterRole called checkout collides across namespaces — include the release name and namespace).

The environment layout question — one namespace per environment in one cluster, or separate clusters — is out of scope here, but the Helm-side rule is simple: the release name should not encode the environment (checkout, not checkout-prod), because the namespace or cluster already does, and encoding it twice means every values file carries a name override.

Try it yourself: run a deliberately failing upgrade twice — once with --atomic and once without — and watch what the running pods do in each case. In both, the old pods keep serving; the difference is whether the release is left in failed for a human to notice. That difference is the whole argument for the flag.

Common mistake: running helm upgrade without --wait in CI, seeing a green pipeline, and telling everyone the deploy succeeded. Helm returned when the API server accepted the manifests, which says nothing about whether a single pod started. The image can be missing, the probe can fail, the quota can reject the pods — and the pipeline is green throughout. --atomic (which implies --wait) makes the pipeline tell the truth.