Rendering a chart correctly is the easy part. The decisions that matter in production are how the chart gets tested before anyone trusts it, and which system is responsible for the cluster matching it afterwards.
Topic 1: Three Delivery Models
CI runs helm upgrade. The pipeline holds cluster credentials and runs the command. Simple, imperative, and drift is invisible until the next deploy. Fine for one cluster and a small team.
GitOps reconciles. An in-cluster controller — Argo CD or Flux — watches Git and makes the cluster match. The pipeline holds no cluster credentials; it updates a file. Drift is detected and corrected continuously, and rollback is a revert.
helm template | kubectl apply. Render to plain YAML and apply it. No release state, so no helm history, no helm rollback, no helm test against a release. Occasionally the right answer — air-gapped delivery, a one-shot install, or a pipeline that already owns its own rollback logic — and a real loss of capability when chosen without noticing.
# Model 1
helm upgrade --install checkout ./chart -n prod -f values/prod.yaml \
--atomic --timeout 8m --set image.tag="${GIT_SHA}"
# Model 3
helm template checkout ./chart -f values/prod.yaml --namespace prod \
| kubectl apply -n prod -f -
Topic 2: The Chart CI Pipeline
A chart is code, and the same rules apply: it should be tested before anyone depends on it.
name: chart-ci
on: [pull_request]
jobs:
test:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
with: { fetch-depth: 0 } # ct needs history to detect changed charts
- uses: azure/setup-helm@v4
- name: Lint
run: helm lint charts/checkout --strict
- name: Render every values shape
run: |
for f in charts/checkout/ci/*-values.yaml; do
echo "== $f"
helm template checkout charts/checkout -f "$f" > /dev/null
done
- name: Validate against real API schemas
run: |
helm template checkout charts/checkout | kubeconform -strict -summary \
-schema-location default \
-schema-location 'https://raw.githubusercontent.com/datreeio/CRDs-catalog/main/{{.Group}}/{{.ResourceKind}}_{{.ResourceAPIVersion}}.json'
- uses: helm/kind-action@v1
- name: Install and test in a real cluster
run: |
helm install checkout charts/checkout --wait --timeout 5m
helm test checkout --logs
helm upgrade checkout charts/checkout --wait --timeout 5m # prove upgrades work too
The parts that catch the most:
ci/*-values.yaml— one file per conditional shape. A chart withifblocks has more than one output, and rendering the default only proves one path.kubeconform— validates rendered manifests against real Kubernetes JSON schemas, including CRDs. Catches misspelled fields thathelm lintnever looks at.- kind, then install, then
helm test— the only step that proves the chart works rather than that it parses. - An upgrade in the same job. Installing works far more often than upgrading; testing only install misses immutable-field problems entirely.
chart-testing(ct lint --check-version-increment) enforces the version bump, which is the rule most often forgotten.
Topic 3: GitOps with Argo CD and Flux
Argo CD — an Application points at a chart and supplies values:
apiVersion: argoproj.io/v1alpha1
kind: Application
metadata:
name: checkout
namespace: argocd
spec:
project: default
source:
repoURL: ghcr.io/acme/charts
chart: checkout
targetRevision: 0.4.1
helm:
valueFiles: [$values/deploy/values/prod.yaml]
destination:
server: https://kubernetes.default.svc
namespace: prod
syncPolicy:
automated: { prune: true, selfHeal: true }
syncOptions: [CreateNamespace=true, ServerSideApply=true]
Flux — a HelmRepository plus a HelmRelease:
apiVersion: helm.toolkit.fluxcd.io/v2
kind: HelmRelease
metadata:
name: checkout
namespace: prod
spec:
interval: 5m
chart:
spec:
chart: checkout
version: "0.4.1"
sourceRef: { kind: HelmRepository, name: acme, namespace: flux-system }
values:
replicaCount: 6
install: { remediation: { retries: 3 } }
upgrade: { remediation: { retries: 3, remediateLastFailure: true } }
Two behaviours worth understanding before adopting either:
Argo CD does not use Helm’s release mechanism. It renders the chart with helm template and applies the result itself. There are no release Secrets, helm list shows nothing, and helm rollback does not apply — Argo’s own history and rollback replace them. Flux, by contrast, drives real Helm releases, so helm list and helm history work normally.
Hook semantics differ. Argo maps some Helm hooks onto its own sync phases and ignores others; Flux runs them as Helm does. A migration hook that works with helm upgrade may behave differently under Argo, which is a common and confusing surprise. Test it in the tool you actually use.
The image tag question. Under GitOps nothing should be --set at deploy time — the tag has to be in Git. The usual answers are a CI step that commits the new tag, or an image automation controller (Flux’s image reflector, or Argo CD Image Updater) that does the commit for you.
Topic 4: Diff Before Apply
Whatever the model, previewing is the habit that prevents the most damage:
helm diff upgrade checkout ./chart -f values/prod.yaml -n prod
helm diff upgrade checkout ./chart -f values/prod.yaml -n prod --detailed-exitcode
--detailed-exitcode returns 2 when there are changes, which lets a pipeline gate on “something actually changed” or require an approval step when the diff is non-empty.
Argo CD and Flux both surface the same diff — Argo in its UI and argocd app diff, Flux via flux diff helmrelease. A release nobody previewed is a change nobody reviewed, regardless of how it is applied.
Topic 5: Rollback in Each Model
| Model | Rollback |
|---|---|
| CI runs helm | helm rollback checkout <rev> — fast, imperative, and now Git and the cluster disagree |
| GitOps | Revert the commit; the controller reconciles. Slower, and the record stays consistent |
template | apply | You own it. Re-render the previous chart version and apply |
The GitOps caveat during an incident: helm rollback on a GitOps-managed release is undone within minutes by the controller, which reconciles back to Git. Under Argo CD or Flux the emergency procedure is either to revert the commit or to suspend reconciliation first:
flux suspend helmrelease checkout -n prod
# ...intervene...
flux resume helmrelease checkout -n prod
Write that into the runbook. Discovering it while a controller fights your fix is a bad way to learn it.
Topic 6: Choosing, and Chart-Level Requirements
Use CI-runs-helm when there is one cluster, a small team, and the pipeline is already the deployment mechanism for everything else.
Use GitOps when there is more than one cluster or environment, when you want drift detection, or when you would rather not give a pipeline cluster credentials.
Use template | apply when something structural prevents the others, and record what you gave up.
Whichever you pick, the chart itself has to hold up:
□ every values shape has a ci/*-values.yaml and renders
□ helm test exercises the real contract, not just DNS
□ the chart version is bumped when chart files change (enforced in CI)
□ dependencies are pinned and Chart.lock is committed
□ no secrets in any values file
□ hooks have before-hook-creation and are GitOps-tested
□ migrations are backward compatible with the previous release
□ helm diff runs before every production upgrade
Try it yourself: deploy the same chart with a pipeline helm upgrade and with an Argo CD Application, then kubectl scale the Deployment in both. One snaps back within minutes; the other stays changed until the next deploy. That difference is what “drift detection” means in practice, and it is the strongest argument for the GitOps column.
Common mistake: running helm upgrade from CI against a release that Argo CD or Flux also manages. Two systems now believe they own the release, and they alternate — your deploy applies, the controller reverts it, the next reconcile flips it back. Pick one owner per release and make it obvious in the repository which one it is.