Debugging Charts

Four failure stages with four different tools — knowing whether you are looking at a values problem, a render problem, an API rejection or a runtime failure.

intermediate 18 min lesson hands-on task included

Helm failures are easy once you know which of four stages you are in. The stages fail differently, and using a stage-1 tool on a stage-4 problem is why debugging sometimes feels like guesswork.


Topic 1: The Four Stages

FOUR STAGES, FOUR DIFFERENT FAILURES — KNOW WHICH ONE YOU ARE IN 1 MERGE VALUES fails as: wrong value wins helm template … --debug 2 RENDER TEMPLATES fails as: Go template error, bad YAML helm template · helm lint 3 SEND TO API SERVER fails as: schema invalid, immutable field, RBAC helm install --dry-run=server 4 RESOURCES RECONCILE fails as: ImagePullBackOff, probe fails, timeout kubectl · helm status THE MOST USEFUL HABIT IN THIS MODULE helm template . -f prod.yaml | kubectl apply --dry-run=server -f - — stages 1–3 checked, nothing changed.
Each row fails in its own way and has its own tool. The command at the bottom checks stages 1 to 3 in one line without changing anything — it belongs in every pipeline and in your shell history.

Stage 1 — values merge. Nothing has rendered yet. The failure is “the wrong value won”, and the tool is --debug, which prints the computed values.

Stage 2 — template render. Go template errors, and YAML that does not parse. Local, fast, no cluster needed: helm template, helm lint.

Stage 3 — API server. The YAML is valid but Kubernetes rejects it — unknown field, invalid value, admission webhook, RBAC, immutable field. Only a server can tell you: --dry-run=server.

Stage 4 — runtime. Everything applied, and the workload does not work. Helm’s job is over; this is kubectl territory, and Helm only appears in it as the messenger reporting a timeout.


Topic 2: Stage 1 and 2 — Local Rendering

# The complete merged values, then every rendered manifest
helm template checkout . -f values/prod.yaml --debug

# One file at a time — turns "somewhere in eleven templates" into one file
helm template checkout . --show-only templates/deployment.yaml

# Validate chart structure, metadata and obvious mistakes
helm lint . --strict

# Render with the values a live release is using
helm get values checkout -n prod -a > /tmp/live.yaml
helm template checkout . -f /tmp/live.yaml

--debug prints the computed values above the manifests. That block answers most “why is it doing that” questions immediately, because it shows the result of the merge rather than the four inputs to it.

Rendered YAML that will not parse is the most common stage-2 failure, and it is almost always indentation. Pipe the output through a parser to get a precise location:

helm template . | kubectl apply --dry-run=client -f - 2>&1 | head
helm template . | yq . > /dev/null      # yq points at the exact line

A useful trick when the error is opaque: render to a file and look at the region by line number.

helm template . > /tmp/out.yaml
sed -n '40,60p' /tmp/out.yaml

Seeing the actual output beside the template usually makes the missing nindent obvious.


Topic 3: Stage 3 — What Only a Server Knows

helm install checkout . -f values/prod.yaml --dry-run=server --debug
helm upgrade checkout . -f values/prod.yaml --dry-run=server
helm template . | kubectl apply --dry-run=server -f -

--dry-run=client (the old --dry-run) only renders. --dry-run=server sends the manifests to the API server for validation and applies nothing, which catches:

  • Unknown or misspelled fields — client-side rendering happily produces containerPorts instead of ports.
  • Invalid values — a CPU limit of "100" where a quantity was expected.
  • Admission webhooks — Pod Security Admission, OPA/Gatekeeper, Kyverno. These reject at admission, so nothing local can predict them.
  • RBAC — you may not be allowed to create what the chart contains.
  • Immutable field changes on an existing object.

That last one makes --dry-run=server particularly valuable before an upgrade: it turns “the deploy failed halfway through” into “this would fail, here is the field”.


Topic 4: Stage 4 — Runtime

When Helm reports timed out waiting for the condition, Helm is not the problem; it waited and something did not become ready.

helm status checkout -n prod

kubectl get all -n prod -l app.kubernetes.io/instance=checkout
kubectl describe deploy/checkout -n prod
kubectl describe pod -n prod -l app.kubernetes.io/instance=checkout | sed -n '/Events:/,$p'
kubectl logs -n prod -l app.kubernetes.io/instance=checkout --tail=50 --all-containers
kubectl get events -n prod --sort-by=.lastTimestamp | tail -20

The usual causes, and what names each one:

SymptomWhere it shows
ImagePullBackOffPod events — wrong tag, wrong registry, missing pull secret
CrashLoopBackOffkubectl logs --previous — the app is exiting
Pending foreverPod events — no capacity, unbound PVC, taints
Ready never trueDescribe the pod — readiness probe path, port or timing
Rollout stuck at N/Mkubectl rollout status — often a PDB or a quota

The label selector is the shortcut. Every well-written chart sets app.kubernetes.io/instance={{ .Release.Name }}, which means one selector scopes every command to one release — a large part of why the standard labels are worth having.


Topic 5: When Helm Itself Is Stuck

Three states that are Helm’s own problem rather than the workload’s.

another operation (install/upgrade/rollback) is in progress — a previous run was interrupted and the release is left in pending-*.

helm history checkout -n prod         # confirm the pending revision

# If the previous revision is healthy, roll back to it — this clears the lock
helm rollback checkout <last-good-revision> -n prod

# If rollback refuses, the last resort is editing the release Secret's status
kubectl get secret -n prod -l owner=helm,name=checkout \
  --sort-by=.metadata.creationTimestamp
# then patch the offending revision's label/status, or delete the pending revision Secret

Deleting the pending revision’s Secret is effective and blunt: it removes Helm’s record of that attempt. Check what is actually running first, because the cluster state does not change when you delete the record of it.

cannot re-use a name that is still in use — a release with that name exists, possibly uninstalled with --keep-history. helm list -a -n prod shows it.

rendered manifests contain a resource that already exists — an object exists that Helm does not own, usually created by hand or by a previous non-Helm deployment. Either delete it, or adopt it by adding the ownership metadata Helm expects:

kubectl annotate deploy/checkout -n prod \
  meta.helm.sh/release-name=checkout meta.helm.sh/release-namespace=prod
kubectl label deploy/checkout -n prod app.kubernetes.io/managed-by=Helm

Adoption is a real technique for migrating hand-managed resources into a chart without downtime, and it is worth knowing before you are tempted to delete production objects.


Topic 6: Drift — Is Production Running the Chart?

Helm does not reconcile. Between upgrades, anything can change an object and Helm will not notice until the next upgrade quietly reverts it.

# What Helm believes it applied vs what the chart renders today
diff <(helm get manifest checkout -n prod) \
     <(helm template checkout . -f values/prod.yaml --namespace prod)

# What Helm applied vs what is actually live
helm diff upgrade checkout . -f values/prod.yaml -n prod     # the plugin

helm diff upgrade is the closest thing Helm has to terraform plan, and it is the single most useful plugin in the ecosystem. Running it before every production upgrade turns “I hope this only changes the image tag” into a reviewed statement.

A debugging checklist worth keeping, in order:

1. helm template --debug          → values merged as expected?
2. helm template --show-only F    → does this file render?
3. --dry-run=server               → would the API server accept it?
4. helm history / helm status     → what does Helm think the state is?
5. kubectl describe + events      → what is the workload actually doing?
6. helm get manifest vs template  → has something drifted?

Try it yourself: break one thing at each stage and time how long each tool takes to name the cause. The pattern that emerges — local tools take a second, cluster tools take ten, and using them in the wrong order costs minutes — is the reason the order matters.

Common mistake: debugging a template error by running helm upgrade repeatedly against a real cluster. Each iteration is slow, each failure may leave the release in pending-upgrade, and the error text mixes template failures with API rejections. Render locally until the YAML is right, validate with --dry-run=server once, and only then touch the cluster.