Jenkins is usually the least reproducible system in an engineering organisation: years of UI clicks, an unknown plugin set, and a backup nobody has restored. Fixing that is unglamorous and it is the difference between an outage and an afternoon.
Topic 1: The Three Surfaces
Configuration — what the controller is: security, agents, tools, libraries, jobs. This belongs in Git.
State — what the controller has done: build history, artifacts, fingerprints, credentials. This needs a backup.
Risk — plugins, script approvals, permissions. This needs a schedule.
The failure mode almost everyone has is using a JENKINS_HOME tarball for all three. It restores yesterday’s controller including yesterday’s undocumented UI changes, and it tells you nothing about what the configuration is supposed to be.
Topic 2: Configuration as Code
JCasC turns the UI into a view of a YAML file:
# jenkins.yaml
jenkins:
systemMessage: "Managed by JCasC — do not configure in the UI"
numExecutors: 0
mode: EXCLUSIVE
quietPeriod: 0
securityRealm:
oic: # or ldap, or github oauth
clientId: "${OIDC_CLIENT_ID}"
clientSecret: "${OIDC_CLIENT_SECRET}"
wellKnownOpenIDConfigurationUrl: "https://sso.acme.example/.well-known/openid-configuration"
authorizationStrategy:
roleBased:
roles:
global:
- name: admin
permissions: ["Overall/Administer"]
entries: [{ group: "platform-team" }]
- name: developer
permissions: ["Overall/Read", "Job/Build", "Job/Cancel"]
entries: [{ group: "engineering" }]
clouds:
- kubernetes:
name: k8s
namespace: jenkins-agents
jenkinsUrl: http://jenkins.jenkins.svc.cluster.local:8080
containerCapStr: "50"
templates:
- name: default
label: k8s
containers:
- name: jnlp
image: jenkins/inbound-agent:latest-jdk17
resourceRequestCpu: "500m"
resourceRequestMemory: "512Mi"
credentials:
system:
domainCredentials:
- credentials:
- string:
id: "npm-publish-token"
secret: "${NPM_TOKEN}" # from the environment, never inline
scope: GLOBAL
unclassified:
globalLibraries:
libraries:
- name: acme-ci
defaultVersion: v3.2.0
implicit: false
retriever:
modernSCM:
scm: { git: { remote: "https://github.com/acme/jenkins-library.git" } }
location:
url: "https://jenkins.acme.example/"
adminAddress: "platform@acme.example"
CASC_JENKINS_CONFIG=/var/jenkins_home/jenkins.yaml
Three practices that make this work:
- Secrets come from the environment,
${VAR}, injected from a real secret store — never written into the YAML. - Reload without a restart from the UI (Manage Jenkins → Configuration as Code → Reload) or by API, so a config change is a fast, reviewable operation.
- Export what already exists to bootstrap the file: the same screen has an Export button, and its output is a starting point rather than a finished file.
Jobs as code complete the picture — a Job DSL seed job, or multibranch organisation folders that discover repositories automatically. The goal is that no job is created by hand.
Topic 3: Plugins
# plugins.txt — pinned, and reviewed like a dependency file
configuration-as-code:1836.v1b_dc_a_c5c9c3d
workflow-aggregator:600.vb_57cdd26fdd7
git:5.4.1
kubernetes:4267.v05f0ba_a_f22b_e
credentials-binding:681.vf91669a_32e45
pipeline-stage-view:2.34
FROM jenkins/jenkins:2.462.3-lts-jdk17
COPY plugins.txt /usr/share/jenkins/ref/plugins.txt
RUN jenkins-plugin-cli -f /usr/share/jenkins/ref/plugins.txt
COPY jenkins.yaml /var/jenkins_home/jenkins.yaml
ENV CASC_JENKINS_CONFIG=/var/jenkins_home/jenkins.yaml
ENV JAVA_OPTS="-Djenkins.install.runSetupWizard=false"
The operating rules:
- Pin versions.
:latestinplugins.txtmeans the controller changes when it is rebuilt, which is the opposite of reproducible. - Upgrade on a schedule, monthly or with security advisories, in a non-production controller first with real jobs.
- Audit annually and remove. Every installed plugin is third-party code with full controller access; the ones nobody uses are pure risk.
- Subscribe to the security advisories. Most Jenkins compromises in the wild are an unpatched plugin, not a Jenkins core bug.
- Prefer a container step over a plugin when there is a choice — it is versioned with the pipeline and has no access to the controller.
// Find plugins nothing uses — a starting point for the annual prune
Jenkins.instance.pluginManager.plugins
.findAll { !it.hasLicensesXml() }
.each { println "${it.shortName} ${it.version}" }
Topic 4: Backup, and the Drill
What to back up, and what to skip:
#!/bin/bash
set -euo pipefail
JH=/var/jenkins_home
OUT=/backup/jenkins-$(date +%F-%H%M).tgz
tar czf "$OUT" -C "$JH" \
config.xml jenkins.yaml credentials.xml secrets users nodes \
--exclude='jobs/*/builds/*/archive' \
--exclude='jobs/*/builds' \
jobs plugins
# Verify the archive is readable before trusting it
tar tzf "$OUT" > /dev/null && echo "ok: $OUT"
secrets/andcredentials.xmlare a pair. Either without the other is useless, and both together are the highest-value secret in the estate — encrypt the backup and restrict who can read it.jobs/*/buildsis the bulk. Decide deliberately whether build history is worth its size; often the last 30 builds per job is plenty.- Excluding
workspace/is always right; it is scratch space.
The drill is the part that matters. A backup nobody has restored is a hypothesis:
Quarterly: restore into a scratch controller, start it, open a job, run a build.
Record the elapsed time and everything that did not come back.
Things that commonly do not come back, and are worth discovering in a drill rather than an incident: agent secrets (agents must be re-registered), plugin state for plugins not in plugins.txt, and anything configured in the UI after the last backup.
Topic 5: Upgrades
1. Read the LTS upgrade guide and the changelog for skipped versions.
2. Upgrade the non-production controller. Run the ten noisiest jobs.
3. Upgrade plugins first, or core first — follow the guide, they interact.
4. Take a backup immediately before, and keep the previous image tag.
5. Upgrade production in a window, with the rollback being the old image
plus the pre-upgrade JENKINS_HOME.
6. Watch the log on startup. Plugin incompatibilities appear there, not in the UI.
The rollback for a containerised controller is the previous image and the pre-upgrade volume snapshot, which is a strong argument for running Jenkins in a container with JENKINS_HOME on a snapshot-capable volume even if everything else about your Jenkins is traditional.
Do not skip LTS lines by more than one or two. Jumping several versions at once turns a routine upgrade into a plugin-compatibility project.
Topic 6: Monitoring the Controller
The metrics that predict problems, via the Prometheus plugin or the Metrics plugin:
| Signal | Why it matters |
|---|---|
| Queue length and queue wait time | The user-visible measure of capacity |
| Executor utilisation | Idle capacity, or none |
| JVM heap usage and GC pause time | A GC-thrashing controller is slow in every direction |
JENKINS_HOME disk free | A full disk corrupts jobs in creative ways |
| Build duration per job, over time | Where the pipeline is slowly rotting |
| Failed-build rate per job | Which pipeline nobody trusts |
Alert on queue wait time and disk free. They are the two that turn into “Jenkins is down” reports, and both are avoidable with a week’s notice.
And log the audit trail. The Audit Trail plugin records who ran what, who changed configuration, and who approved a script — the record you will want after any incident involving the build system.
Try it yourself: delete your non-production controller and rebuild it from jenkins.yaml plus plugins.txt. Whatever you had to do by hand afterwards is the gap in your configuration-as-code, and the list is usually short and specific.
Common mistake: adopting JCasC but continuing to make changes in the UI “just this once”. The file and the running controller diverge, the next reload silently reverts somebody’s fix, and the team concludes JCasC is unreliable. Put the system message at the top of the config — managed by JCasC, do not configure in the UI — and mean it.