Cloud Build: The Build Configuration

Every field of a build step, substitutions and secrets done safely, the artifacts block, and the four caching strategies — with the measurement that tells you which one helped.

intermediate 24 min lesson hands-on task included

Cloud Build’s configuration is small enough to learn completely, which is unusual and worth exploiting: there is no plugin system to search, no DSL to master. There are steps, and steps are containers.


Topic 1: The Model, in One Picture

CLOUD BUILD — EVERY STEP IS A CONTAINER, SHARING ONE /workspace cloudbuild.yaml steps: - name: gcr.io/cloud-builders/docker args: ['build','-t','$_IMAGE:$SHORT_SHA','.'] - name: gcr.io/cloud-builders/docker args: ['push','$_IMAGE:$SHORT_SHA'] substitutions: { _IMAGE: … } $PROJECT_ID · $SHORT_SHA · $BRANCH_NAME are built in HOW IT DIFFERS FROM JENKINS · no controller to run or patch · each step is an image, not a plugin · identity is a service account, not a stored key · billed per build-minute · steps run in order unless you set waitFor · logs land in Cloud Logging by default · no UI to click a job together — that is the point trigger push, PR, tag, Pub/Sub, manual worker pool default, or PRIVATE in your VPC /workspace shared by every step artifacts Artifact Registry + provenance THE TWO THINGS THAT CATCH PEOPLE OUT 1. Caching is not automatic — mount a volume or pull the previous image with --cache-from. PRIVATE POOLS ARE THE ANSWER TO "IT CANNOT REACH" The default pool has no route into your VPC. A private pool does — and it is how builds reach internal services.
Steps are containers sharing one /workspace, run in order unless you say otherwise. The right-hand panel is the trade against Jenkins, and the bottom-right is the answer to the most common Cloud Build support question.

Before the field-by-field detail: a build is a list of container executions against a shared workspace, started by a trigger, running in a worker pool, as a service account, producing artifacts. Everything in this lesson is one of those five nouns.


Topic 2: The Step Schema, in Full

ONE STEP, EVERY FIELD — THIS IS THE WHOLE SCHEMA name the container image to run. There is no plugin system. gcr.io/cloud-builders/docker args arguments to the image entrypoint ['build','-t','$_IMAGE','.'] entrypoint override the image entrypoint bash script inline shell, instead of entrypoint + args #!/bin/bash … id a name, so other steps can wait for it build waitFor ['-'] = start now · ['id'] = after that step ['-'] dir working directory inside /workspace services/api env plain environment variables ['GOOS=linux'] secretEnv from availableSecrets — never in args ['API_KEY'] volumes persist a path between steps name: go-mod timeout per step, separate from the build timeout 600s allowFailure do not fail the build if this step fails true allowFailure and allowExitCodes turn a gate into a report — use them deliberately, not to make a red build green.
Twelve fields, and you will use eight of them regularly. The two at the bottom are the ones that quietly turn a gate into a report.
steps:
  - name: golang:1.23                 # the image — this IS the extensibility model
    id: test                          # a name other steps can waitFor
    entrypoint: bash                  # override the image entrypoint
    dir: services/api                 # working directory inside /workspace
    env:
      - 'GOFLAGS=-mod=readonly'
      - 'CGO_ENABLED=0'
    secretEnv: ['CODECOV_TOKEN']      # from availableSecrets — never in args
    timeout: 600s                     # per step
    args:
      - -c
      - |
        set -euo pipefail
        go test ./... -coverprofile=cover.out

The fields that repay a second look:

script is the modern alternative to entrypoint: bash plus args: ['-c', …], and it reads better:

  - name: golang:1.23
    script: |
      #!/usr/bin/env bash
      set -euo pipefail
      go build ./...

waitFor turns the step list into a DAG. By default each step waits for the previous one; waitFor: ['-'] means “no dependencies, start immediately”, and naming step IDs expresses anything in between:

  - id: build   { … }
  - id: lint    { waitFor: ['-'] }          # runs alongside build
  - id: test    { waitFor: ['build'] }
  - id: package { waitFor: ['test', 'lint'] }

allowFailure: true and allowExitCodes: [1, 2] let a step fail without failing the build. They have narrow legitimate uses — an optional report, a best-effort cache warm — and they are also the easiest way to turn a security gate into a decoration. If you use them, say why in a comment.

volumes persist a path between steps, which is how you keep a module cache within one build:

  - name: golang:1.23
    volumes: [{ name: gomod, path: /go/pkg/mod }]
    script: go mod download
  - name: golang:1.23
    volumes: [{ name: gomod, path: /go/pkg/mod }]     # same volume, second step
    script: go build ./...

/workspace is already shared between all steps; volumes is for paths outside it.


Topic 3: Substitutions, and the Ones That Bite

Built-in, always present:

$PROJECT_ID   $PROJECT_NUMBER   $BUILD_ID   $LOCATION
$COMMIT_SHA   $SHORT_SHA        $REVISION_ID
$BRANCH_NAME  $TAG_NAME         $REF_NAME   $REPO_NAME   $REPO_FULL_NAME

$SHORT_SHA and $COMMIT_SHA are empty for manual builds submitted with gcloud builds submit unless you pass them. A config that works from a trigger and produces image: (with nothing after the colon) by hand is this, every time:

gcloud builds submit --substitutions=SHORT_SHA=$(git rev-parse --short HEAD)

User substitutions start with an underscore:

substitutions:
  _IMAGE: europe-west1-docker.pkg.dev/${PROJECT_ID}/apps/checkout
  _REGION: europe-west1
  _SERVICE: checkout

options:
  dynamicSubstitutions: true        # lets one substitution reference another
  substitutionOption: MUST_MATCH    # the default — fail on an unused/missing one

Keep MUST_MATCH. ALLOW_LOOSE silently ignores a typo in a substitution name, and the result is an empty string in the middle of an image reference.

Bash-style defaults work inside values with dynamicSubstitutions:

substitutions:
  _TAG: ${SHORT_SHA:-manual}

Topic 4: Secrets

SECRETS COME FROM SECRET MANAGER, AS ENV VARS, PER STEP Secret Manager versions, IAM, audit log availableSecrets declared once per build secretEnv named per step that needs it availableSecrets: secretManager: - versionName: projects/$PROJECT_ID/secrets/npm-token/versions/latest env: NPM_TOKEN steps: [{ name: node, entrypoint: sh, args: ['-c','npm publish'], secretEnv: ['NPM_TOKEN'] }] NEVER PUT A SECRET IN args args are recorded in the build metadata and visible to anyone with builds.get. secretEnv is not. gcloud builds describe BUILD_ID — read your own args back THE BUILD IDENTITY IS THE REAL CONTROL The service account needs secretmanager.secretAccessor on that secret — scope it per trigger, not per project. Every access appears in the Secret Manager audit log.
Declared once, named per step. The left panel is the reason `args` is the wrong place for a secret — build metadata is readable by anyone who can describe the build.
availableSecrets:
  secretManager:
    - versionName: projects/$PROJECT_ID/secrets/npm-token/versions/latest
      env: NPM_TOKEN
    - versionName: projects/$PROJECT_ID/secrets/signing-key/versions/3
      env: SIGNING_KEY

steps:
  - name: node:22
    entrypoint: sh
    secretEnv: ['NPM_TOKEN']
    args:
      - -c
      - |
        set -euo pipefail
        echo "//registry.npmjs.org/:_authToken=${NPM_TOKEN}" > ~/.npmrc
        npm publish

Four rules, in order of how badly each is missed:

  • Never put a secret in args or env. Both are stored in the build resource and returned by gcloud builds describe to anyone with cloudbuild.builds.get. secretEnv is not.
  • The service account needs secretmanager.secretAccessor on that specific secret — grant it per secret, not project-wide.
  • Pin the version (/versions/3) for anything where a rotation should be a deliberate change; use latest where automatic pickup is the point.
  • Every access is in the Secret Manager audit log, which is a materially better trail than a CI system’s own store.

For a file rather than an environment variable, write it inside the step and never persist it to /workspace beyond the step that needs it.


Topic 5: The artifacts Block

Pushing with a docker push step works. Declaring artifacts does more:

artifacts:
  images:
    - '${_IMAGE}:${SHORT_SHA}'
    - '${_IMAGE}:latest'
  objects:
    location: 'gs://acme-releases/${_SERVICE}/${SHORT_SHA}/'
    paths: ['dist/*.tar.gz', 'sbom.json']
  mavenArtifacts:
    - repository: 'https://europe-west1-maven.pkg.dev/acme/maven'
      path: 'target/checkout-1.0.jar'
      artifactId: 'checkout'
      groupId: 'example.acme'
      version: '${SHORT_SHA}'

What declaring images: gives you beyond the push:

  • The digest is recorded on the build, so gcloud builds describe tells you exactly what was produced — the value you should be passing downstream.
  • Build provenance is generated, which is what Binary Authorization can verify later. The supply-chain lesson uses it.
  • The push happens after all steps succeed, so a failing test cannot leave a half-published tag behind.
# The digest, for whatever deploys next
gcloud builds describe "$BUILD_ID" \
  --format='value(results.images[0].digest)'

Topic 6: Caching, Measured

EVERY BUILD STARTS ON A CLEAN WORKER — CACHING IS ALWAYS EXPLICIT docker --cache-from pull the previous image reuse matching layers needs an extra pull step Dockerfile builds, simplest to a dd kaniko --cache layer cache in a registry no daemon needed --cache-ttl bounds staleness The lowest-effort default GCS rsync ~/.m2, ~/.npm, go mod copy in, copy out shared mutable state Dependency caches volumes persist between STEPS not between builds zero cost Passing data step to step MEASURE BEFORE AND AFTER A cache that saves 20 seconds and costs 15 in pull time is not a cache. Record both build durations. gcloud builds list --format='table(id,duration)' A SHARED CACHE IS SHARED MUTABLE STATE A poisoned dependency cache survives retries and follows the workload. Key it by lockfile hash, and know how to clear it.
Four strategies with different scopes. The panel on the left is the discipline that matters — a cache you have not measured may be costing you time.

Kaniko’s registry cache — usually the least effort:

  - name: gcr.io/kaniko-project/executor:latest
    args:
      - --context=.
      - --destination=${_IMAGE}:${SHORT_SHA}
      - --cache=true
      - --cache-repo=${_IMAGE}/cache
      - --cache-ttl=168h

Docker --cache-from — one extra pull step:

  - name: gcr.io/cloud-builders/docker
    entrypoint: bash
    args:
      - -c
      - |
        docker pull ${_IMAGE}:latest || true
        docker build --cache-from ${_IMAGE}:latest \
          -t ${_IMAGE}:${SHORT_SHA} -t ${_IMAGE}:latest .

A dependency cache in GCS — for package managers:

  - id: cache-restore
    name: gcr.io/cloud-builders/gsutil
    waitFor: ['-']
    args: ['-m', 'rsync', '-r', 'gs://acme-cache/npm', '/workspace/.npm']
  # …build…
  - id: cache-save
    name: gcr.io/cloud-builders/gsutil
    args: ['-m', 'rsync', '-r', '/workspace/.npm', 'gs://acme-cache/npm']

Then measure, because a cache that costs more to fetch than it saves is common:

gcloud builds list --limit=10 \
  --format='table(id, status, createTime.date("%H:%M"), duration)'

Record cold, warm-with-layer-cache and warm-with-dependency-cache. The dependency cache in particular is often a wash for small projects and transformative for large ones — and you cannot tell which without the numbers.

The caution from the agents lesson applies here too: a shared cache is shared mutable state. Key it by lockfile hash where you can, and know the command to clear it before you need it at 6pm.


Topic 7: A Complete, Defensible Config

steps:
  - id: deps
    name: node:22
    entrypoint: sh
    args: ['-c', 'npm ci --cache /workspace/.npm --prefer-offline']

  - id: lint
    name: node:22
    waitFor: ['deps']
    entrypoint: sh
    args: ['-c', 'npm run lint']

  - id: test
    name: node:22
    waitFor: ['deps']
    entrypoint: sh
    args: ['-c', 'npm test -- --ci --reporters=default --reporters=jest-junit']

  - id: image
    name: gcr.io/kaniko-project/executor:latest
    waitFor: ['lint', 'test']
    args:
      - --context=.
      - --destination=${_IMAGE}:${SHORT_SHA}
      - --cache=true
      - --digest-file=/workspace/digest

  - id: sbom
    name: anchore/syft:latest
    entrypoint: sh
    args: ['-c', 'syft ${_IMAGE}@$(cat /workspace/digest) -o spdx-json > /workspace/sbom.json']

  - id: scan
    name: anchore/grype:latest
    entrypoint: sh
    args: ['-c', 'grype sbom:/workspace/sbom.json --fail-on high --only-fixed']

substitutions:
  _IMAGE: europe-west1-docker.pkg.dev/${PROJECT_ID}/apps/checkout

artifacts:
  images: ['${_IMAGE}:${SHORT_SHA}']
  objects:
    location: 'gs://acme-build-artifacts/${SHORT_SHA}/'
    paths: ['/workspace/sbom.json']

options:
  machineType: E2_HIGHCPU_8
  logging: CLOUD_LOGGING_ONLY
  substitutionOption: MUST_MATCH
  dynamicSubstitutions: true

timeout: 1200s

Read it against the pipeline contract from lesson 1: one build, a digest captured and carried forward, lint and test in parallel, a scan that fails rather than warns, and the artifact published declaratively. That is roughly the whole of what a build stage owes you.

Try it yourself: run the config, then run gcloud builds describe on the build ID and read the whole resource. Every substitution, every arg, the images and their digests are in there — which is both the audit trail and the reason secrets must never be in args.

Common mistake: using latest as the only tag and passing that tag downstream. The build is reproducible, the deployment is not: latest moves, and what you scanned is not necessarily what you deployed. Capture the digest with --digest-file or from the build result, and pass that.