Agents and Build Environments

Static agents versus ephemeral pods, why a pinned container per stage beats installing tools on a VM, and building images without handing out the Docker socket.

intermediate 20 min lesson hands-on task included

Where a build runs decides how reproducible it is, how fast it starts, and how much damage a malicious dependency can do. This is the lesson where CI stops being about Jenkins and starts being about containers.


Topic 1: The Four Kinds of Agent

Static agent — a long-lived VM with tools installed. Fast to start, and it drifts: one agent has Java 17, another has 21, and the build that fails on linux-03 becomes folklore. Every build shares the filesystem with the last one.

Docker agent — the stage runs inside a container on an agent host. The toolchain is pinned per stage in the Jenkinsfile, and the container is discarded.

Kubernetes pod agent — a pod per build, created on demand and deleted afterwards. The standard for Jenkins on Kubernetes.

Cloud agent — an EC2 or GCE instance provisioned per build. Useful when a build needs a whole machine, and slower to start.

// Static
agent { label 'linux && jdk21' }

// Docker
agent {
    docker {
        image 'maven:3.9-eclipse-temurin-21'
        args '-v $HOME/.m2:/root/.m2'
        label 'docker'
    }
}

// Dockerfile from the repository — the build environment is versioned too
agent {
    dockerfile {
        filename 'ci/Dockerfile'
        additionalBuildArgs '--build-arg JDK=21'
    }
}

// Kubernetes
agent {
    kubernetes {
        yamlFile 'ci/agent-pod.yaml'
        defaultContainer 'maven'
    }
}

Label expressions support &&, || and !, which is how you express real requirements: linux && !gpu, arm64 || amd64.


Topic 2: Kubernetes Pod Agents

ONE POD PER BUILD — CREATED ON DEMAND, DELETED AFTERWARDS 1 build queued no agent exists yet 2 pod created jnlp + your tool containers 3 steps run container("maven") { … } 4 pod deleted workspace gone with it THE POD IS A MULTI-CONTAINER SPEC containers: - name: jnlp # talks to the controller - name: maven # your build tool - name: kaniko # builds images, no docker socket They share the workspace volume, so a step in one sees files from another. WHAT THIS FIXES · no snowflake agents drifting apart · no leftover files between unrelated builds · capacity scales with the cluster · tool versions are declared per stage And it moves your CI capacity problem into Kubernetes. DO NOT MOUNT THE DOCKER SOCKET TO BUILD IMAGES /var/run/docker.sock in a build pod is root on the node. Use kaniko, buildkit rootless, or a build service — all covered in the container module.
One pod per build, deleted afterwards. The panel at the bottom is the one rule to carry away — a mounted Docker socket in a build pod is root on the node.
# ci/agent-pod.yaml
apiVersion: v1
kind: Pod
spec:
  securityContext:
    runAsNonRoot: true
    runAsUser: 1000
  containers:
    - name: maven
      image: maven:3.9-eclipse-temurin-21
      command: ['sleep']
      args: ['infinity']
      resources:
        requests: { cpu: '1', memory: 2Gi }
        limits:   { memory: 4Gi }
      volumeMounts:
        - { name: m2, mountPath: /root/.m2 }
    - name: kaniko
      image: gcr.io/kaniko-project/executor:v1.23.2-debug
      command: ['sleep']
      args: ['infinity']
  volumes:
    - name: m2
      persistentVolumeClaim: { claimName: maven-cache }
pipeline {
    agent { kubernetes { yamlFile 'ci/agent-pod.yaml'; defaultContainer 'maven' } }
    stages {
        stage('Build') {
            steps { sh 'mvn -B package' }        // runs in the maven container
        }
        stage('Image') {
            steps {
                container('kaniko') {
                    sh '''/kaniko/executor --context=. --dockerfile=Dockerfile \
                          --destination=$IMAGE:$GIT_COMMIT --cache=true'''
                }
            }
        }
    }
}

Details that decide whether this works well:

  • Every container needs sleep infinity (or equivalent) as its command. Jenkins runs steps by exec-ing into a container that must already be running.
  • The jnlp container is injected automatically. Override it only to pin its image or give it resources.
  • Containers share the workspace volume, so a file written in maven is visible in kaniko.
  • Set requests and limits. Without them the pod lands in the BestEffort QoS class and is the first thing evicted under pressure — which presents as random build failures.
  • Pod startup is a real cost — image pull plus scheduling, typically 10–40 seconds. Cache images on nodes, keep agent images small, and consider a small pool of idle agents for interactive work.

Topic 3: Building Images Without the Docker Socket

The tempting pattern is to mount /var/run/docker.sock into the build so docker build works. That grants root on the node — the container can start a privileged container, mount the host filesystem, and read every secret on that machine. On a shared build cluster it is equivalent to handing out cluster-admin.

The alternatives, all fine in practice:

ToolNotes
kanikoBuilds from a Dockerfile in userspace, no daemon. The common default.
BuildKit rootlessbuildctl against a rootless buildkitd; supports advanced cache mounts.
BuildahRootless, good for OCI-native workflows.
Cloud Build / a build serviceThe build happens outside your cluster entirely — the Cloud Build lesson covers this.
// buildkit, rootless, with a shared cache
container('buildkit') {
    sh '''
      buildctl-daemonless.sh build \
        --frontend dockerfile.v0 --local context=. --local dockerfile=. \
        --output type=image,name=$IMAGE:$GIT_COMMIT,push=true \
        --export-cache type=registry,ref=$IMAGE:buildcache,mode=max \
        --import-cache type=registry,ref=$IMAGE:buildcache
    '''
}

The container module covers image layers and cache behaviour; the point here is that a CI system never needs the Docker socket, and treating that as a hard rule removes the single largest privilege escalation in most build platforms.


Topic 4: Caches, and Where They Bite

Caching is where most CI time is won, and where most cross-build contamination comes from.

// A PVC shared across builds — fast, and shared state
volumeMounts: [{ name: 'm2', mountPath: '/root/.m2' }]

// Registry-backed layer cache — no shared filesystem
--export-cache type=registry,ref=$IMAGE:buildcache

The trade-off, stated plainly: a shared cache volume is fast and is shared mutable state between builds. A poisoned or corrupted cache produces failures that survive a retry and follow the workload around. Two mitigations that keep the speed:

  • Key the cache by lockfile hash where the tool supports it, so a dependency change gets a fresh cache rather than a mutated one.
  • Prefer read-through caches — a registry pull-through mirror, or a remote build cache — over a read-write volume, because they cannot be corrupted by one bad build.

And a diagnostic worth knowing: when a build fails only on one agent, or only after another team’s build ran, suspect the cache before the code.


Topic 5: Sizing and Concurrency

options {
    disableConcurrentBuilds(abortPrevious: true)   // per branch
    lock(resource: 'staging-environment')          // exclusive access to something shared
    throttleJobProperty(categories: ['heavy-builds'], maxConcurrentTotal: 4)
}

lock (from the Lockable Resources plugin) is the answer to “two pipelines deployed to staging at once and interleaved”. It is a semaphore across the controller, and it is far more reliable than a convention.

disableConcurrentBuilds(abortPrevious: true) is right for most branch builds: when three commits land in a minute, only the newest matters.

For capacity, the honest guidance is to measure queue time. If builds wait more than a minute for an executor during normal hours, add agents; if they never wait, you are paying for idle capacity. On Kubernetes this becomes a cluster autoscaling question, and the Kubernetes and cost modules cover it from that side.


Topic 6: Reproducibility

The property worth aiming at: the same commit produces the same artifact, whoever builds it and whenever.

□ tool versions pinned in the agent image, not installed at build time
□ the agent image itself pinned by digest, not by a moving tag
□ dependencies pinned by a committed lockfile
□ base images pinned by digest in the Dockerfile
□ no `latest` anywhere in the pipeline
□ build timestamps and paths normalised where the toolchain allows

Full bit-for-bit reproducibility is hard and rarely required. Pinning is cheap and gets most of the value: it turns “the build broke and nothing changed” into a statement that is actually true, because when nothing changed, nothing changed.

Try it yourself: run a build on a Docker agent pinning maven:3.9-eclipse-temurin-21, then change the tag to maven:latest and run it again a week later. The first is reproducible; the second is a build whose result depends on the date.

Common mistake: installing tools on static agents with a configuration management script and treating that as the build environment. Agents drift, the script is not run everywhere, and a build’s toolchain is now a property of the host rather than of the commit. Pin the environment in the Jenkinsfile — a container image per stage — and the build becomes portable between agents, clouds and eventually CI systems.