Debugging Pipelines and Flaky Builds

A triage order for a failing build, why replay is the fastest loop you have, and what to do about a test suite nobody believes any more.

advanced 20 min lesson hands-on task included

A pipeline that fails for a reason nobody can explain is worse than no pipeline, because it trains a team to ignore red. This lesson is about diagnosing quickly and about the specific discipline that flakiness requires.


Topic 1: The Triage Order

FAILING BUILD — ANSWER THESE IN ORDER, NOT BY INTUITION 1 Did it fail, or did it never start? queue blocked, no agent with that label, quota 2 Same commit, different result? that is FLAKY — not a code failure 3 Did the pipeline change, or the code? git log on the Jenkinsfile and the library tag 4 Is it the environment? agent disk full, cache poisoned, registry down 5 Reproduce it locally in the same image docker run the agent image, run the step 6 Only then, read the application logs by which point you usually know the answer Replay (pipeline replay) re-runs a build with an edited script — the fastest loop for step 3, and it never touches the repo.
Answer these in order. Step 2 is the one that changes the entire investigation — a build that fails intermittently on an unchanged commit is not a code failure and should not be debugged as one.
1. Did it fail, or did it never start?
   Queue blocked · no agent with that label · cloud quota · plugin error on startup

2. Same commit, different result?
   → FLAKY. Stop looking at the code. Go to Topic 4.

3. Did the pipeline change, or the code?
   git log on the Jenkinsfile · the shared library tag · the agent image digest

4. Is it the environment?
   agent disk full · poisoned cache · registry down · a certificate expired

5. Reproduce it locally, in the same image
   docker run the agent image and run the failing step by hand

6. Only now, read the application logs

Steps 1–4 take a minute each and resolve most failures. The instinct to start at step 6 is what turns a ten-minute diagnosis into an afternoon.

The single most useful check is step 3, because “nothing changed” is almost never true: the agent image moved, the library tag was updated, a base image was rebuilt, or a dependency resolved differently.


Topic 2: Replay — the Fast Loop

Editing a Jenkinsfile, committing, pushing and waiting is a two-minute cycle for a one-character fix. Replay re-runs a completed build with an edited script, without touching the repository:

Build page → Replay → edit the pipeline → Run

It applies to the Jenkinsfile and to loaded shared-library code, which makes it the fastest way to test a library change. When you have it working, commit the version Replay proved.

Two limits worth knowing: Replay needs the Job/Replay permission (often not granted to everyone by default), and it runs the edited script against the original commit — so it cannot test a change that depends on new repository content.

The other diagnostic tools, in the order they help:

// Print what the pipeline actually thinks it knows
sh 'env | sort'
echo "node=${env.NODE_NAME} ws=${env.WORKSPACE} branch=${env.BRANCH_NAME}"

// Keep the workspace on failure so you can look at it
post { failure { archiveArtifacts artifacts: '**/target/*.log', allowEmptyArchive: true } }

// Pause a running build to inspect the agent
input message: 'Paused for inspection — check the workspace, then continue'

That last trick — an input step in a temporary Replay — gives you a live agent with the exact build state, which beats reasoning about it from logs.


Topic 3: Reading the Failure Correctly

SymptomWhere to look
Cannot run program … No such fileThe tool is not in the agent image. Pin the image, do not install at build time
Build hangs, then the timeout firesA process waiting on input, a network call with no timeout, a lock
NotSerializableExceptionA non-serialisable object held across a step — usually in a library class
MissingPropertyException: No such propertyA typo, or script { } scope confusion
Works on one agent, fails on anotherAgent drift, or a cache left by a previous build
Passes locally, fails in CIEnvironment: user, HOME, TZ, locale, DNS, or a missing service
ERROR: script returned exit code 1 and nothing elseThe tool printed to stderr and it was swallowed — add set -o pipefail
Queued foreverNo agent matches the label, or the cloud cannot provision

The pipefail one is worth internalising. A sh step running make | tee build.log reports success when make fails, because the pipeline’s exit code is tee’s. Every non-trivial shell step should start with:

sh '''
  set -euo pipefail
  make | tee build.log
'''

Topic 4: Flaky Tests

A flaky test is one that passes and fails on the same commit. It is not a test failure; it is a defect in the test suite, and the response is different.

First, measure it. Without a number, this is an argument about impressions:

# Ten runs of the same commit — the pass rate is the deliverable
for i in $(seq 1 10); do
  ./gradlew test --tests 'com.acme.CheckoutTest' && echo PASS || echo FAIL
done | sort | uniq -c

Then quarantine, do not retry. Retrying makes the suite slower and hides the signal:

stage('Test') {
    steps {
        sh './gradlew test -PexcludeTags=flaky'
    }
}
stage('Flaky (non-blocking)') {
    steps {
        catchError(buildResult: 'SUCCESS', stageResult: 'UNSTABLE') {
            sh './gradlew test -PincludeTags=flaky'
        }
    }
}

The quarantined tests still run and still report, and they no longer block anyone. The rule that keeps quarantine honest: a quarantined test has an owner and a deadline, and a quarantine list that only grows is a suite being abandoned slowly.

The usual causes, in rough order of frequency: shared mutable state between tests, timing assumptions (sleep instead of waiting for a condition), test-order dependence, real network calls, and unbounded concurrency in the test runner. All five are fixable; none are fixed by retry(3).


Topic 5: Durability and Restarts

Jenkins serialises pipeline state at each step so a build can survive a controller restart. That is why library classes must be Serializable, and it has two operational levers:

options {
    // Trade durability for speed on high-volume jobs
    durabilityHint('PERFORMANCE_OPTIMIZED')
}

MAX_SURVIVABILITY (the default) writes state frequently and is slower; PERFORMANCE_OPTIMIZED writes less and a controller restart loses in-flight builds. For a busy CI job that will simply be re-triggered, the second is a reasonable trade; for a long deploy pipeline, it is not.

Restart from a stage is a declarative-only feature and is genuinely useful: when a 20-minute build fails at the deploy step for an environmental reason, you can restart from Deploy without rebuilding. It requires that earlier stages’ outputs are recoverable — which in practice means artifacts must be stashed or published rather than left in a workspace that may be gone.


Topic 6: Making Failures Cheap

The properties that make a pipeline pleasant to debug are mostly decisions made earlier:

□ stages named after what they achieve, so the stage view localises failures
□ every sh step starts with set -euo pipefail
□ test results published in post { always }, not only on success
□ logs and reports archived on failure
□ agent images pinned by digest, so "nothing changed" is verifiable
□ the shared library pinned to a tag, for the same reason
□ timeouts everywhere, so a hang is a failure rather than a mystery
□ one lock around anything shared, so interleaving is impossible
□ build duration and outcome exported as metrics, so rot is visible

And track the flake rate as a first-class metric. A dashboard showing “percentage of builds that passed on retry” tells you whether the pipeline is trusted, which is the thing that actually determines whether people ship small changes often.

Try it yourself: run the same commit ten times and count the failures. Teams routinely believe their suite is reliable and discover a 15% flake rate — at which point one in seven honest failures is being ignored as noise.

Common mistake: responding to intermittent failures by adding retries at every level — retry the test, retry the stage, retry the build. The pipeline goes green, the underlying defect stays, build times triple on bad days, and a real regression now looks exactly like the noise everyone has learned to re-run. Measure the flake rate, quarantine with an owner, and keep the signal.