Jenkins Operations: JCasC, Plugins and Backup

Rebuilding a controller from a file in Git instead of a backup, keeping the plugin surface patched, and proving the restore works before you need it.

advanced 20 min lesson hands-on task included

Jenkins is usually the least reproducible system in an engineering organisation: years of UI clicks, an unknown plugin set, and a backup nobody has restored. Fixing that is unglamorous and it is the difference between an outage and an afternoon.


Topic 1: The Three Surfaces

THE THREE OPERATIONAL SURFACES OF A JENKINS CONTROLLER CONFIGURATION jenkins.yaml (JCasC) plugins.txt, pinned job DSL or multibranch seed job in Git Rebuild the controller from Git, not from a STATE jobs/*/builds — history secrets/ + credentials.xml fingerprints, artifacts plugin data Back it up, and restore it once to prove it RISK plugins — the CVE surface script approvals agent-to-controller access anonymous read Patch weekly. Most Jenkins incidents are a p JCasC IN ONE LINE CASC_JENKINS_CONFIG=/var/jenkins_home/casc.yaml — the UI becomes a view of a file you review in a pull request. The test: delete the controller. If you cannot rebuild it in an hour from Git plus one backup, that is the gap to close first.
Configuration should come from Git, state from a backup, and risk from a patch schedule. Most Jenkins pain comes from trying to make the backup cover all three.

Configuration — what the controller is: security, agents, tools, libraries, jobs. This belongs in Git.

State — what the controller has done: build history, artifacts, fingerprints, credentials. This needs a backup.

Risk — plugins, script approvals, permissions. This needs a schedule.

The failure mode almost everyone has is using a JENKINS_HOME tarball for all three. It restores yesterday’s controller including yesterday’s undocumented UI changes, and it tells you nothing about what the configuration is supposed to be.


Topic 2: Configuration as Code

JCasC turns the UI into a view of a YAML file:

# jenkins.yaml
jenkins:
  systemMessage: "Managed by JCasC — do not configure in the UI"
  numExecutors: 0
  mode: EXCLUSIVE
  quietPeriod: 0

  securityRealm:
    oic:                                  # or ldap, or github oauth
      clientId: "${OIDC_CLIENT_ID}"
      clientSecret: "${OIDC_CLIENT_SECRET}"
      wellKnownOpenIDConfigurationUrl: "https://sso.acme.example/.well-known/openid-configuration"

  authorizationStrategy:
    roleBased:
      roles:
        global:
          - name: admin
            permissions: ["Overall/Administer"]
            entries: [{ group: "platform-team" }]
          - name: developer
            permissions: ["Overall/Read", "Job/Build", "Job/Cancel"]
            entries: [{ group: "engineering" }]

  clouds:
    - kubernetes:
        name: k8s
        namespace: jenkins-agents
        jenkinsUrl: http://jenkins.jenkins.svc.cluster.local:8080
        containerCapStr: "50"
        templates:
          - name: default
            label: k8s
            containers:
              - name: jnlp
                image: jenkins/inbound-agent:latest-jdk17
                resourceRequestCpu: "500m"
                resourceRequestMemory: "512Mi"

credentials:
  system:
    domainCredentials:
      - credentials:
          - string:
              id: "npm-publish-token"
              secret: "${NPM_TOKEN}"      # from the environment, never inline
              scope: GLOBAL

unclassified:
  globalLibraries:
    libraries:
      - name: acme-ci
        defaultVersion: v3.2.0
        implicit: false
        retriever:
          modernSCM:
            scm: { git: { remote: "https://github.com/acme/jenkins-library.git" } }
  location:
    url: "https://jenkins.acme.example/"
    adminAddress: "platform@acme.example"
CASC_JENKINS_CONFIG=/var/jenkins_home/jenkins.yaml

Three practices that make this work:

  • Secrets come from the environment, ${VAR}, injected from a real secret store — never written into the YAML.
  • Reload without a restart from the UI (Manage Jenkins → Configuration as Code → Reload) or by API, so a config change is a fast, reviewable operation.
  • Export what already exists to bootstrap the file: the same screen has an Export button, and its output is a starting point rather than a finished file.

Jobs as code complete the picture — a Job DSL seed job, or multibranch organisation folders that discover repositories automatically. The goal is that no job is created by hand.


Topic 3: Plugins

# plugins.txt — pinned, and reviewed like a dependency file
configuration-as-code:1836.v1b_dc_a_c5c9c3d
workflow-aggregator:600.vb_57cdd26fdd7
git:5.4.1
kubernetes:4267.v05f0ba_a_f22b_e
credentials-binding:681.vf91669a_32e45
pipeline-stage-view:2.34
FROM jenkins/jenkins:2.462.3-lts-jdk17
COPY plugins.txt /usr/share/jenkins/ref/plugins.txt
RUN jenkins-plugin-cli -f /usr/share/jenkins/ref/plugins.txt
COPY jenkins.yaml /var/jenkins_home/jenkins.yaml
ENV CASC_JENKINS_CONFIG=/var/jenkins_home/jenkins.yaml
ENV JAVA_OPTS="-Djenkins.install.runSetupWizard=false"

The operating rules:

  • Pin versions. :latest in plugins.txt means the controller changes when it is rebuilt, which is the opposite of reproducible.
  • Upgrade on a schedule, monthly or with security advisories, in a non-production controller first with real jobs.
  • Audit annually and remove. Every installed plugin is third-party code with full controller access; the ones nobody uses are pure risk.
  • Subscribe to the security advisories. Most Jenkins compromises in the wild are an unpatched plugin, not a Jenkins core bug.
  • Prefer a container step over a plugin when there is a choice — it is versioned with the pipeline and has no access to the controller.
// Find plugins nothing uses — a starting point for the annual prune
Jenkins.instance.pluginManager.plugins
    .findAll { !it.hasLicensesXml() }
    .each { println "${it.shortName} ${it.version}" }

Topic 4: Backup, and the Drill

What to back up, and what to skip:

#!/bin/bash
set -euo pipefail
JH=/var/jenkins_home
OUT=/backup/jenkins-$(date +%F-%H%M).tgz

tar czf "$OUT" -C "$JH" \
  config.xml jenkins.yaml credentials.xml secrets users nodes \
  --exclude='jobs/*/builds/*/archive' \
  --exclude='jobs/*/builds' \
  jobs plugins

# Verify the archive is readable before trusting it
tar tzf "$OUT" > /dev/null && echo "ok: $OUT"
  • secrets/ and credentials.xml are a pair. Either without the other is useless, and both together are the highest-value secret in the estate — encrypt the backup and restrict who can read it.
  • jobs/*/builds is the bulk. Decide deliberately whether build history is worth its size; often the last 30 builds per job is plenty.
  • Excluding workspace/ is always right; it is scratch space.

The drill is the part that matters. A backup nobody has restored is a hypothesis:

Quarterly:  restore into a scratch controller, start it, open a job, run a build.
            Record the elapsed time and everything that did not come back.

Things that commonly do not come back, and are worth discovering in a drill rather than an incident: agent secrets (agents must be re-registered), plugin state for plugins not in plugins.txt, and anything configured in the UI after the last backup.


Topic 5: Upgrades

1. Read the LTS upgrade guide and the changelog for skipped versions.
2. Upgrade the non-production controller. Run the ten noisiest jobs.
3. Upgrade plugins first, or core first — follow the guide, they interact.
4. Take a backup immediately before, and keep the previous image tag.
5. Upgrade production in a window, with the rollback being the old image
   plus the pre-upgrade JENKINS_HOME.
6. Watch the log on startup. Plugin incompatibilities appear there, not in the UI.

The rollback for a containerised controller is the previous image and the pre-upgrade volume snapshot, which is a strong argument for running Jenkins in a container with JENKINS_HOME on a snapshot-capable volume even if everything else about your Jenkins is traditional.

Do not skip LTS lines by more than one or two. Jumping several versions at once turns a routine upgrade into a plugin-compatibility project.


Topic 6: Monitoring the Controller

The metrics that predict problems, via the Prometheus plugin or the Metrics plugin:

SignalWhy it matters
Queue length and queue wait timeThe user-visible measure of capacity
Executor utilisationIdle capacity, or none
JVM heap usage and GC pause timeA GC-thrashing controller is slow in every direction
JENKINS_HOME disk freeA full disk corrupts jobs in creative ways
Build duration per job, over timeWhere the pipeline is slowly rotting
Failed-build rate per jobWhich pipeline nobody trusts

Alert on queue wait time and disk free. They are the two that turn into “Jenkins is down” reports, and both are avoidable with a week’s notice.

And log the audit trail. The Audit Trail plugin records who ran what, who changed configuration, and who approved a script — the record you will want after any incident involving the build system.

Try it yourself: delete your non-production controller and rebuild it from jenkins.yaml plus plugins.txt. Whatever you had to do by hand afterwards is the gap in your configuration-as-code, and the list is usually short and specific.

Common mistake: adopting JCasC but continuing to make changes in the UI “just this once”. The file and the running controller diverge, the next reload silently reverts somebody’s fix, and the team concludes JCasC is unreliable. Put the system message at the top of the config — managed by JCasC, do not configure in the UI — and mean it.