Project: Package, Publish and Operate a Chart

Build a chart somebody else could operate — schema, subchart, hooks, tests, signed and published — then put it through nine drills including the ones that fail.

advanced 45 min lesson hands-on task included

Everything in this module, assembled once, for a service you actually run. The chart is half the deliverable; the drills are the other half, and they are what tell you whether the chart is operable rather than merely renderable.


Topic 1: The Target

PROJECT TARGET — A CHART SOMEBODY ELSE COULD OPERATE THE CHART MUST HAVE □ values.yaml that documents every knob□ values.schema.json rejecting bad input□ _helpers.tpl with name/label templates□ resources, probes, PDB, HPA — all optional□ a subchart, enabled by a condition□ a pre-upgrade hook with a delete policy□ templates/tests/ that curls the service□ NOTES.txt with the real next command□ a signed .tgz pushed to an OCI registry AND SURVIVE THESE DRILLS 1. install into a clean namespace2. upgrade with a changed image tag3. upgrade that FAILS a probe → --atomic4. rollback to the previous revision5. an immutable-field change, handled6. a hook failure, diagnosed from its pod7. helm test passes against the release8. install the same chart twice, one namespace9. uninstall, leaving no orphans behind THE TEST OF A GOOD CHART: SOMEONE ELSE CAN CONFIGURE IT WITHOUT READING templates/ If a reviewer must open a template to find out what a value does, the values file and the schema are the things to fix. DELIVERABLE The chart, a README documenting every value, a CI job that lints/templates/installs into kind and runs helm test, and a drill log with timings.
The left column is the chart; the right column is what it must survive. The line in the middle is the acceptance test — if a reviewer has to open templates/ to configure it, the values file and the schema are what need work.

Pick a real service — one you maintain, or a small application you write for this. It needs at minimum a Deployment, a Service, configuration, and one dependency.

The requirements, stated as constraints:

1. Installs twice into one namespace under different release names,
   with no collisions.
2. Every configurable knob is in values.yaml and documented.
3. values.schema.json rejects three classes of bad input.
4. One subchart, switchable with a condition.
5. A pre-upgrade hook that is safe to roll back past.
6. helm test verifies the service actually answers.
7. Published to an OCI registry, signed, and installable by version.
8. A CI pipeline that would have caught every mistake you made building it.

Constraint 1 is the one that catches sloppy templating, and constraint 5 is the one that catches sloppy migration design.


Topic 2: Phase 1 — The Chart

charts/checkout/
├── Chart.yaml              version + appVersion moving independently
├── values.yaml             every knob, every one commented
├── values.schema.json      types, enums, required keys
├── README.md               generated by helm-docs from values.yaml
├── templates/
│   ├── _helpers.tpl        name, fullname, labels, selectorLabels, image
│   ├── deployment.yaml     probes, resources, securityContext, checksum/config
│   ├── service.yaml
│   ├── ingress.yaml        optional, behind ingress.enabled
│   ├── configmap.yaml
│   ├── hpa.yaml            optional — and replicas NOT templated when enabled
│   ├── pdb.yaml            optional
│   ├── serviceaccount.yaml
│   ├── migrate-job.yaml    pre-upgrade hook, weight -5, before-hook-creation
│   ├── NOTES.txt           the real next command
│   └── tests/
│       └── smoke.yaml      curls the health endpoint AND one real route
├── ci/
│   ├── default-values.yaml
│   ├── ingress-values.yaml
│   ├── autoscaling-values.yaml
│   └── minimal-values.yaml
└── .helmignore

Decisions to write down as you make them, because these are what a reviewer will ask about:

  • Why each value is a value rather than a hardcoded field — and what you deliberately did not expose.
  • Which labels are in selectorLabels and why nothing version-bearing is there.
  • Whether the chart templates replicas, and what happens when autoscaling.enabled is true.
  • What the pre-upgrade hook does, and what a rollback past it leaves behind.
  • Which registry the images come from, and how the tag defaults.

Verify before moving on:

helm lint charts/checkout --strict
for f in charts/checkout/ci/*-values.yaml; do helm template checkout charts/checkout -f "$f" >/dev/null || echo "FAILED $f"; done
helm template checkout charts/checkout | kubectl apply --dry-run=server -f -

Topic 3: Phase 2 — Dependency, Schema and Tests

The subchart. Add a dependency with a condition — a database for local development that production disables in favour of a managed one:

dependencies:
  - name: postgresql
    version: "15.5.x"
    repository: https://charts.bitnami.com/bitnami
    condition: postgresql.enabled

Prove both paths render, and that the application’s connection string is correct in each — pointing at the subchart’s service when enabled, at an external host when not. That conditional connection string is where most umbrella charts have a bug.

The schema. Make it reject three genuinely different things:

{
  "required": ["image"],
  "properties": {
    "replicaCount": { "type": "integer", "minimum": 0 },
    "image": { "type": "object", "required": ["repository"] },
    "service": { "properties": { "type": { "enum": ["ClusterIP", "NodePort", "LoadBalancer"] } } }
  }
}

Then demonstrate each rejection: a missing required key, a wrong type, and a value outside an enum.

The test. Not the generated one:

args:
  - |
    set -e
    URL="http://{{ include "checkout.fullname" . }}:{{ .Values.service.port }}"
    [ "$(curl -s -o /dev/null -w '%{http_code}' "$URL/healthz")" = "200" ]
    curl -sf "$URL/api/v1/status" | grep -q '"database":"ok"'

Topic 4: Phase 3 — The Nine Drills

Each drill: establish the state, run it, verify with a command, record the time.

1. Clean install. helm install into an empty namespace with --wait. Verify with helm test.

2. Ordinary upgrade. Change the image tag. Verify the new tag is running: kubectl get deploy -o jsonpath='{..image}'.

3. Failing upgrade with --atomic. Point the readiness probe at a 404. Expect: the upgrade fails, rolls back automatically, and the old pods never stopped serving. Record: how long the whole cycle took, and confirm helm history shows both the failure and the rollback.

4. Manual rollback. Upgrade twice, then helm rollback to revision 2. Verify the running image matches revision 2 and that a new revision 4 was created.

5. An immutable-field change. Add a label to selectorLabels and upgrade. Expect: rejection. Recover deliberately with --cascade=orphan and record the downtime, if any. Then revert the chart change and explain why selectorLabels must stay invariant.

6. A hook failure. Make the migration Job exit 1. Expect: the release fails before the workload changes. Read the failed hook’s logs — which requires that you left hook-failed out of the delete policy. Fix and re-run.

7. Two releases, one namespace. helm install checkout and helm install checkout-canary from the same chart. Expect: no collisions. Any failure here is a hardcoded name.

8. Config change rolls pods. Change a ConfigMap value and upgrade. Expect: pods restart, because of the checksum/config annotation. Remove the annotation, repeat, and observe that they do not — that contrast is the point.

9. Clean uninstall. helm uninstall, then search for orphans:

kubectl get all,pvc,secret,cm -n <ns> -l app.kubernetes.io/instance=checkout

Anything left is either a bug or a deliberate resource-policy: keep, and your notes should say which.


Topic 5: Phase 4 — Publish and Automate

helm dependency build charts/checkout
helm package charts/checkout --sign --key 'release@acme.example' --destination dist/
helm verify dist/checkout-0.4.1.tgz

helm registry login ghcr.io -u "$USER" --password-stdin <<< "$TOKEN"
helm push dist/checkout-0.4.1.tgz oci://ghcr.io/acme/charts

# Install from the published artifact, in a clean cluster, as a consumer would
helm install checkout oci://ghcr.io/acme/charts/checkout --version 0.4.1 --wait
helm test checkout

The pipeline, which should fail on every mistake you made by hand while building this:

on pull_request:   lint --strict · render every ci/ values · kubeconform ·
                   kind install · helm test · upgrade in place ·
                   fail if the chart version was not bumped
on tag:            dependency build · package --sign · push to OCI ·
                   install from the registry into kind as a final check

Topic 6: What to Produce, and How It Is Judged

Four artifacts:

  1. The chart repository, with the CI pipeline described above and a ci/ directory covering every conditional path.
  2. A generated README (helm-docs) documenting every value, its default and its purpose.
  3. A drill log — nine drills, each with the command, the observed result, the verification command and the elapsed time.
  4. An operations note, one page: what a rollback of this chart does and does not undo, which fields are immutable, what the hook changes, and what helm uninstall leaves behind.

The acceptance test: hand the chart to a colleague and ask them to configure it for a new environment. If they can do it from values.yaml, the schema and the README without opening templates/, the chart is finished. If they cannot, the values file is the thing to fix — not the templates.

The questions you should be able to answer without notes:

  • Which of your values are safe to change at 3am, and which need a maintenance window?
  • What happens if someone runs kubectl edit on your Deployment?
  • Which fields in your chart cannot be changed by an upgrade?
  • What does your pre-upgrade hook leave behind after a rollback?
  • Why does helm uninstall not delete the StatefulSet’s PVCs?
  • Which of your ci/ values files covers the path most likely to break in production?

Common mistake: treating the drills as a formality after the chart “works”. Drills 3, 5, 6 and 8 all fail on charts that install perfectly — the atomic rollback, the immutable selector, the hook evidence and the config checksum are exactly the properties that only appear under failure. The install proves the chart renders; the drills prove somebody can operate it.