Install is the easy half. Everything interesting about Helm happens on the second deploy, when there is existing state to reconcile and a decision to make about what happens if the new version does not come up.
Topic 1: Revisions Are Append-Only
helm install checkout ./chart -n prod # revision 1
helm upgrade checkout ./chart -n prod # revision 2
helm upgrade checkout ./chart -n prod # revision 3 (fails)
helm rollback checkout 2 -n prod # revision 4 == revision 2's content
helm history checkout -n prod
REVISION UPDATED STATUS CHART APP VERSION DESCRIPTION
1 Mon Aug 18 08:10 superseded checkout-0.3.9 1.4.0 Install complete
2 Mon Aug 18 08:40 superseded checkout-0.4.0 1.5.0 Upgrade complete
3 Mon Aug 18 09:12 failed checkout-0.4.1 1.6.0 Upgrade "checkout" failed: timed out
4 Mon Aug 18 09:20 deployed checkout-0.4.0 1.5.0 Rollback to 2
The status column is the useful part:
| Status | Meaning |
|---|---|
deployed | The current, live revision |
superseded | A former revision, replaced by a later one |
failed | The operation did not complete |
pending-install / pending-upgrade / pending-rollback | An operation is running — or crashed partway |
uninstalling | Deletion in progress |
A stuck pending-* status is the single most common Helm operational problem. It means a previous command was interrupted — CI timed out, a laptop closed, a pod running Helm was evicted — and Helm now refuses to start another operation with another operation (install/upgrade/rollback) is in progress. The recovery is in the upgrade-failures lesson; recognising the cause here is what makes it a two-minute fix.
Bound the history, or every upgrade grows etcd:
helm upgrade --install checkout ./chart -n prod --history-max 10
Topic 2: The Flags That Matter
helm upgrade --install checkout ./chart -n prod \
-f values-prod.yaml \
--atomic \
--timeout 5m \
--history-max 10
--install — install if absent, upgrade if present. Makes the command idempotent, which is what a pipeline needs.
--wait — do not return until all resources report ready: Deployments to their expected replica count, PVCs bound, Services with endpoints, Jobs completed. Without it, helm upgrade returns as soon as the API server accepts the objects, and a green pipeline means nothing about whether the release works.
--timeout (default 5m) — how long --wait waits. Set it above your slowest legitimate rollout; a chart with a database migration hook and a slow image pull needs more than five minutes, and a timeout that fires on a healthy deploy is worse than no timeout.
--atomic — implies --wait, and rolls back automatically if the upgrade fails or times out. This is the flag that turns a failed deploy into a non-event. The cost: the rollback takes time too, so a failing deploy occupies the pipeline for up to twice the timeout.
--cleanup-on-fail — delete resources created during a failed upgrade instead of leaving them orphaned. Worth adding whenever a chart creates new objects between versions.
--dry-run=server — render, send to the API server for validation, apply nothing. Catches schema errors, admission webhook rejections and RBAC problems that client-side rendering cannot.
--force — deletes and recreates resources rather than patching. It is not a stronger --wait; it causes downtime and can destroy things. There is a narrow legitimate use (an immutable field change you have decided to accept downtime for) and it should never be a reflex in a pipeline.
A defensible production command:
helm upgrade --install checkout ./chart \
--namespace prod --create-namespace \
-f values/base.yaml -f values/prod.yaml \
--set image.tag="${GIT_SHA}" \
--atomic --timeout 8m --history-max 10 \
--description "deploy ${GIT_SHA} by ${CI_JOB_URL}"
--description is underused: it lands in the helm history DESCRIPTION column, which turns the release log into something an on-call engineer can correlate with a pipeline run.
Topic 3: What Rollback Actually Does
helm rollback re-applies a previous revision’s stored manifest. That is a precise and limited promise, and knowing its limits prevents a bad assumption during an incident.
helm rollback checkout # to the previous revision
helm rollback checkout 2 # to a specific one
helm rollback checkout 2 --wait --timeout 5m
It reverts: the Kubernetes objects the release manages — image tags, replica counts, environment variables, resource limits, config.
It does not revert:
- Anything a hook did. A migration Job that ran during the failed upgrade has already changed your database. Helm has no undo for it.
- Data written by the new version. Rows, files, queue messages, cache entries in a new format.
- Deleted PVCs, and it will not shrink a volume that was expanded.
- CRDs, which Helm never upgrades or deletes at all.
- Anything outside the release — a DNS record, an external load balancer created by a controller, a row in another system.
The consequence for schema migrations is the most important operational rule in this module: migrations must be backward compatible with the previous application version, or you do not have a rollback. Add a column, deploy code that writes both, remove the old column two releases later. If the migration drops a column the previous version reads, rolling back the application produces a broken application, faster.
Topic 4: Uninstall, and What Survives It
helm uninstall checkout -n prod
helm uninstall checkout -n prod --keep-history # release record stays, resources go
helm list -n prod --uninstalled
What Helm does not remove:
- PersistentVolumeClaims created by a StatefulSet’s
volumeClaimTemplates. Kubernetes owns those, deliberately, so your data survives. They must be deleted by hand — and they keep costing money until somebody does. - CRDs from
crds/, and therefore every custom resource of those kinds. - Anything with
helm.sh/resource-policy: keep, which is a per-resource opt-out from deletion. - Resources created by hooks, unless a delete policy removed them.
metadata:
annotations:
helm.sh/resource-policy: keep
That annotation is the right tool for a PVC or a Secret that must outlive the release, and a trap when applied carelessly — a kept resource blocks a later reinstall with “already exists”.
After any uninstall, check for orphans:
kubectl get pvc,secret,cm -n prod -l app.kubernetes.io/instance=checkout
Topic 5: Inspecting a Live Release
When something is wrong, these four commands answer most questions before you open the chart:
helm status checkout -n prod # current revision, resources, NOTES
helm get values checkout -n prod # what values were supplied
helm get values checkout -n prod -a # ...merged with chart defaults (the real config)
helm get manifest checkout -n prod # exactly what was applied
helm get hooks checkout -n prod # hook manifests attached to this release
helm get manifest is the ground truth. It is what Helm actually sent to the API server for the current revision — not what the chart in your working directory renders today. The difference between the two is drift, and comparing them is the fastest way to answer “is production running what is in main?”:
diff <(helm get manifest checkout -n prod) \
<(helm template checkout ./chart -f values/prod.yaml --namespace prod)
helm get values -a deserves its own mention: without -a you see only the overrides, which hides the defaults that are doing the real work. During an incident, -a is the flag you want.
Topic 6: Deploying the Same Chart Repeatedly
A chart is a package; a release is an installation. Several releases of one chart in one namespace is normal and is a good test of chart quality:
helm install checkout ./chart -n prod
helm install checkout-canary ./chart -n prod -f values/canary.yaml
For this to work, every resource name must derive from .Release.Name, every selector must include app.kubernetes.io/instance, and nothing may use a cluster-scoped fixed name (a ClusterRole called checkout collides across namespaces — include the release name and namespace).
The environment layout question — one namespace per environment in one cluster, or separate clusters — is out of scope here, but the Helm-side rule is simple: the release name should not encode the environment (checkout, not checkout-prod), because the namespace or cluster already does, and encoding it twice means every values file carries a name override.
Try it yourself: run a deliberately failing upgrade twice — once with --atomic and once without — and watch what the running pods do in each case. In both, the old pods keep serving; the difference is whether the release is left in failed for a human to notice. That difference is the whole argument for the flag.
Common mistake: running helm upgrade without --wait in CI, seeing a green pipeline, and telling everyone the deploy succeeded. Helm returned when the API server accepted the manifests, which says nothing about whether a single pod started. The image can be missing, the probe can fail, the quota can reject the pods — and the pipeline is green throughout. --atomic (which implies --wait) makes the pipeline tell the truth.