Almost every team says they have CI/CD. Far fewer can say what happens between a merge and production without opening a wiki page that is out of date. This lesson is about the distinctions that matter and the numbers that settle the argument.
Topic 1: Three Terms, One Real Difference
Continuous integration — every developer merges to trunk frequently, and every merge is built and tested automatically. The point is not the build server; it is the frequently. A team where branches live three weeks has a build server and no continuous integration, because the integration problems are all still ahead of them.
Continuous delivery — every green build is releasable. It has been through the full pipeline and is sitting one button-press away from production. A human decides when.
Continuous deployment — no button. Every green build reaches production automatically.
The step from delivery to deployment is smaller than it sounds technically and much larger organisationally: it requires tests you trust, progressive rollout, and monitoring that catches what tests do not. Most teams should get delivery right first, and many should stop there deliberately.
The measurable definition of CI, which cuts through the arguing: if a developer’s work is not merged to trunk at least daily, you are not doing continuous integration whatever the tool says.
Topic 2: The Pipeline Is a Contract
A pipeline makes a series of promises about anything that reaches production:
- It came from a specific commit, and you can name it.
- It was built once. The bytes in production are the bytes that were tested.
- It passed a defined set of checks, and you can list them.
- Someone or something authorised it, and that decision is recorded.
- It can be undone without a rebuild.
Promise 2 is the one that is quietly broken most often. A pipeline that runs docker build in the deploy job for each environment produces three different images with three different digests, from three different sets of transitive dependencies resolved at three different times. What staging tested is not what production runs, and the difference appears exactly when a base image or a package registry changed in between.
Build once, promote the artifact, change only the configuration. That single rule eliminates an entire class of “it worked in staging”.
Topic 3: What Belongs in a Pipeline, and in What Order
Order by how fast a check fails and how much it costs to run:
seconds lint, format, compile, unit tests
→ fail here and nobody waits
1–5 minutes integration tests, container build, SBOM + vulnerability scan
→ the bulk of your signal
5–20 minutes end-to-end tests, performance smoke, security scans
→ run on merge, not necessarily on every push
on demand full load test, chaos experiments, manual exploratory testing
Two rules that keep the pipeline useful:
Fail fast and fail loudly. A ten-minute pipeline that fails at minute nine on a lint error is a design mistake. Put the cheap checks first.
Every check either blocks or is deleted. A scan that reports and proceeds is a report nobody reads. If a finding should not stop the build, it should not be in the build — put it in a dashboard.
Topic 4: The Four Numbers
The DORA metrics are the closest thing to an objective answer about a delivery system:
| Metric | What it measures | Rough elite band |
|---|---|---|
| Deployment frequency | How often you ship to production | On demand, multiple per day |
| Lead time for changes | Commit → running in production | Under an hour |
| Change failure rate | Deployments causing a degradation | Under 15% |
| Time to restore | Degradation → recovered | Under an hour |
Two things about them that matter more than the bands:
They resist gaming in pairs. Frequency alone rewards shipping carelessly; frequency plus change failure rate does not. Lead time alone rewards skipping tests; lead time plus time-to-restore does not.
Time to restore is the one to fix first if you have to choose. A system that recovers in five minutes can afford to take risks; one that takes four hours cannot, and every other improvement is bounded by that.
# Deployment frequency and lead time, from Git alone, if you tag releases
git log --tags --simplify-by-decoration --pretty='%ci %d' | head -20
# Lead time for one change: commit time → deploy time
git show -s --format=%ci <sha>
Start with a spreadsheet and honest numbers rather than a dashboard and estimated ones.
Topic 5: Pipeline as Code, and Why It Is Not Optional
A pipeline configured through a web UI has no diff, no review, no history, and no way to answer “what changed about the build last Tuesday”. A pipeline defined in a file in the repository has all four, and it changes in the same pull request as the code it builds.
That principle is why this module treats the Jenkinsfile, the cloudbuild.yaml and the Cloud Deploy pipeline definition as the actual subject, and the web interface as a view of them.
The properties to insist on, whatever the tool:
- The definition lives in the repository it builds, so a branch can change its own pipeline and have that reviewed.
- It is deterministic — pinned tool versions, pinned base images, pinned library versions. A build whose result depends on the day it ran is not reproducible.
- It runs the same way locally as far as possible, so debugging does not require a push.
- Nothing important is configured only in the server UI, because that is the part that disappears when the server does.
Topic 6: The Failure Modes to Recognise Early
Four patterns that show up in almost every mature-but-unhappy pipeline:
The pipeline nobody trusts. Flaky tests mean a red build is ignored and re-run until green. At that point the pipeline is theatre, and the fix is to quarantine flaky tests aggressively rather than to add retries everywhere.
The pipeline that takes 45 minutes. Developers batch changes to avoid waiting, batches get larger, failures get harder to diagnose, and lead time compounds. Measure per-stage duration and attack the top item.
The snowflake deploy. Production is deployed by a different mechanism than staging — a script, a person, a runbook. Everything the pipeline proved about staging is now untested for production.
Secrets everywhere. Every job has access to production credentials because it was easier. One compromised dependency in one test job is then a production incident.
Each of these gets a lesson later in this module. Recognising them now is what makes the rest of the path feel like it is answering questions you already have.
Try it yourself: time your pipeline stage by stage and write the numbers down. Teams are consistently wrong about which stage is slowest, and the measurement usually points at something nobody has looked at in a year.
Common mistake: treating CI/CD as a tooling project. Installing Jenkins, or moving to Cloud Build, changes nothing about branch lifetime, test reliability or deployment size — and those three determine most of the outcome. The tool is the easy half; this module covers it, and keeps saying which half you are in.