Hooks and Chart Tests

Where each hook runs, why a rollback cannot undo one, and how ten lines of chart test turn a green deploy into a verified one.

intermediate 20 min lesson hands-on task included

Hooks let a chart run something at a specific point in the release lifecycle — a migration, a seed job, a smoke check. They are also the part of Helm most likely to break a rollback, so the mechanics and the constraints matter equally.


Topic 1: When Each Hook Runs

HOOKS RUN TO COMPLETION BEFORE HELM CONTINUES — A FAILED HOOK STOPS THE RELEASE pre-install create a schema, seed a bucket the release all normal manifests applied post-install smoke check, register, notify test helm test, on demand THE ANNOTATIONS THAT CONTROL THEM annotations: "helm.sh/hook": pre-upgrade,pre-install "helm.sh/hook-weight": "-5" "helm.sh/hook-delete-policy": before-hook-creation,hook-succeeded READING THOSE THREE LINES weight: lower runs first; ties broken by name before-hook-creation: delete the previous one first hook-succeeded: clean up after a pass omit hook-failed on purpose — keep failures to debug HOOKS ARE NOT PART OF THE RELEASE A rollback does NOT undo what a hook did. A migration hook that ran is permanent — write migrations that are backward compatible, or you have no rollback at all. helm get hooks RELEASE — see what is actually attached helm test — THE CHEAPEST GATE YOU CAN ADD templates/tests/connection.yaml "helm.sh/hook": test A pod that curls your own service and exits non-zero if the answer is wrong. Run it in CI after every install.
Hooks run to completion before Helm continues, and a failure stops the release. The bottom-left panel is the constraint that shapes every migration strategy: rollback does not undo what a hook already did.

A hook is an ordinary manifest with an annotation:

apiVersion: batch/v1
kind: Job
metadata:
  name: {{ include "checkout.fullname" . }}-migrate
  annotations:
    "helm.sh/hook": pre-upgrade,pre-install
    "helm.sh/hook-weight": "-5"
    "helm.sh/hook-delete-policy": before-hook-creation
spec:
  backoffLimit: 0
  template:
    spec:
      restartPolicy: Never
      containers:
        - name: migrate
          image: {{ include "checkout.image" . }}
          command: ["/app/migrate", "up"]

The hook points, in order of use:

HookRuns
pre-installBefore any chart resource is created
post-installAfter all resources are created
pre-upgradeBefore any resource is updated
post-upgradeAfter all resources are updated
pre-rollback / post-rollbackAround a rollback
pre-delete / post-deleteAround an uninstall
testOnly when helm test runs

Helm waits for each hook to reach a completed state before continuing. For a Job, that means Completed; for a Pod, Succeeded. A hook that never completes blocks the release until --timeout fires.

hook-weight orders hooks within the same phase — lower runs first, ties broken by resource name. Weights are strings, and negative values are normal:

"helm.sh/hook-weight": "-10"   # a ConfigMap the migration Job needs
"helm.sh/hook-weight": "-5"    # the migration itself

Topic 2: Delete Policies, and Keeping Failures Around

"helm.sh/hook-delete-policy": before-hook-creation,hook-succeeded
PolicyEffect
before-hook-creationDelete the previous instance before creating this one (the default since Helm 3)
hook-succeededDelete after it succeeds
hook-failedDelete after it fails

before-hook-creation is essential for Jobs, because a Job’s spec.template is immutable. Without it, the second upgrade tries to update the existing Job, is rejected, and the release fails with an error about an immutable field that has nothing to do with your change.

Deliberately omit hook-failed. A failed migration’s pod holds the logs that explain the failure; deleting it removes your only evidence at exactly the moment you need it. before-hook-creation cleans it up on the next attempt anyway.

Hooks are not part of the release. helm get manifest does not show them (helm get hooks does), and Helm does not track them for rollback or deletion the way it tracks normal resources. That is the mechanism behind the constraint in the next topic.


Topic 3: Migrations and the Rollback Constraint

helm rollback restores the previous manifests. It does not undo what a hook did — the migration Job has already altered your database, and Helm has no record of how to reverse it.

Which means: your migrations must be backward compatible with the previous application version, or you do not have a rollback.

The expand-and-contract pattern, which is the only reliable way to keep one:

Release N     add the new column, nullable. Old code ignores it.
Release N+1   write to both old and new. Backfill.
Release N+2   read from new only.
Release N+3   drop the old column.

At every step, rolling back one release leaves a schema the previous application version can still use. A migration that renames a column in one step is a decision to give up rollback for that release, and it should be made consciously rather than discovered at 2am.

Two further considerations for migration hooks:

  • backoffLimit: 0 so a failed migration fails once and fast, rather than retrying six times and burning your --timeout.
  • Concurrency. Two pipelines upgrading the same release simultaneously run two migration Jobs. The migration tool needs its own lock — most have one; verify rather than assume.

A pre-upgrade hook runs before the new pods exist, so the schema change lands while the old version is still serving. That is exactly why backward compatibility is required, and it is also why “add column” is safe and “drop column” is not.


Topic 4: Chart Tests

templates/tests/ holds pods annotated as test hooks. They run only on helm test, never during install.

apiVersion: v1
kind: Pod
metadata:
  name: "{{ include "checkout.fullname" . }}-test"
  annotations:
    "helm.sh/hook": test
    "helm.sh/hook-delete-policy": before-hook-creation
spec:
  restartPolicy: Never
  containers:
    - name: test
      image: curlimages/curl:8.8.0
      command: ["/bin/sh", "-c"]
      args:
        - |
          set -e
          URL="http://{{ include "checkout.fullname" . }}:{{ .Values.service.port }}"
          code=$(curl -s -o /dev/null -w '%{http_code}' "$URL/healthz")
          [ "$code" = "200" ] || { echo "healthz returned $code"; exit 1; }
          curl -sf "$URL/api/v1/status" | grep -q '"database":"ok"'
helm test checkout -n prod
helm test checkout -n prod --logs        # print the test pod's output

The generated test from helm create only checks that the service name resolves, which proves almost nothing. A useful test exercises the actual contract: the health endpoint returns 200, the database connection works, a representative request returns the expected shape.

Where it pays off: in CI, helm install --wait && helm test is the difference between “the chart installed” and “the chart works”. It catches a wrong service port, a missing environment variable, a broken image and a misconfigured dependency — the failures that pass helm lint and --dry-run happily.


Topic 5: Hook Failure Modes

SymptomCause
Job has reached the specified backoff limitThe migration failed. kubectl logs job/<name>-migrate
Upgrade hangs, then times outHook never completes — a Job with no exit, a wrong image
field is immutable on the second upgradeMissing before-hook-creation on a Job
Hook ran but the release still failedHooks pass or fail independently of the workload’s readiness
Hook pod is gone before you can read ithook-failed in the delete policy
Hook ran on install when you wanted upgrade onlyCheck the hook list — pre-install,pre-upgrade runs on both
helm get hooks checkout -n prod                 # what is attached to this release
kubectl get jobs,pods -n prod -l app.kubernetes.io/instance=checkout
kubectl logs -n prod job/checkout-migrate

.Release.IsInstall and .Release.IsUpgrade let one hook behave differently in each case — seeding data on first install only, for example:

{{- if .Release.IsInstall }}
command: ["/app/seed"]
{{- else }}
command: ["/app/migrate", "up"]
{{- end }}

Topic 6: When Not to Use a Hook

Hooks are imperative steps inside a declarative system, and that friction shows up in three places: they are invisible to helm diff, they are not reconciled by GitOps tooling, and they cannot be rolled back.

Alternatives worth considering first:

  • An init container for anything the pod itself needs before starting — waiting for a dependency, fetching config. It is retried automatically and is visible in the pod’s own status.
  • A Kubernetes Job as a normal chart resource, if it does not need to run before the workload.
  • A CronJob for recurring work, which a hook cannot express.
  • The application doing it at startup, with a lock, for simple idempotent migrations.

Under GitOps, hooks need extra care: Argo CD has its own sync-phase annotations and only partially maps Helm hooks; Flux runs them as part of the release. Test the behaviour in your own tooling rather than assuming the semantics carry over — this is a common source of “it worked with helm upgrade but not through Argo”.

Try it yourself: make a pre-upgrade hook exit 1 and run an upgrade. The release fails, the old pods keep serving, and the hook pod is still there with its logs. Then add hook-failed to the delete policy and repeat — the evidence disappears. That contrast is the argument for leaving it out.

Common mistake: putting a database migration in a post-install hook and assuming rollback covers it. It does not, and it never has. The migration ran, the schema is new, and rolling the application back gives you old code against a new schema — which fails in whatever way your ORM fails, usually loudly and in production. Design migrations to be backward compatible and the problem never arises.