Cloud Build: Triggers, Pools and Operations

What starts a build and what it is allowed to do, reaching private networks from a managed builder, and the identity, notification and cost model you inherit.

intermediate 22 min lesson hands-on task included

The build config says what runs. This lesson is everything around it: what starts a build, what identity it carries, what network it sits in, and what it costs — which is where the operational decisions actually are.


Topic 1: Triggers

WHAT STARTS A BUILD, AND WHAT IT IS ALLOWED TO DO push to branch ^main$ build, push, create a release push of a tag ^v[0-9]+\. production release pull request from a fork build and test ONLY — no deploy identity manual gcloud builds triggers run ad-hoc, e.g. a rebuild Pub/Sub a topic message chained builds, upstream artifacts webhook an external system anything outside your forge FILTERS THAT SAVE REAL MONEY includedFiles: ['services/api/**'] ignoredFiles: ['**/*.md', 'docs/**'] In a monorepo this is the difference between one build and twelve. THE FORK PULL-REQUEST BOUNDARY A fork PR runs unreviewed code — including its own cloudbuild.yaml. Require a comment to build, and give that trigger a service account with no deploy rights. Approval on a trigger is a separate, second gate.
Six ways a build starts, and the right-hand column is the part to design deliberately. The bottom-right panel is the boundary that has caused real compromises in every CI ecosystem.
# Push to a branch
gcloud builds triggers create github \
  --name=checkout-main \
  --repo-owner=acme --repo-name=checkout \
  --branch-pattern='^main$' \
  --build-config=cloudbuild.yaml \
  --included-files='services/checkout/**' \
  --ignored-files='**/*.md,docs/**' \
  --service-account=projects/acme/serviceAccounts/cb-checkout-main@acme.iam.gserviceaccount.com \
  --substitutions=_ENVIRONMENT=staging

# Tag push — the production path
gcloud builds triggers create github \
  --name=checkout-release \
  --tag-pattern='^v[0-9]+\.[0-9]+\.[0-9]+$' \
  --build-config=cloudbuild-release.yaml \
  --service-account=projects/acme/serviceAccounts/cb-checkout-release@acme.iam.gserviceaccount.com

# Pull request — build and test only
gcloud builds triggers create github \
  --name=checkout-pr \
  --pull-request-pattern='^main$' \
  --comment-control=COMMENTS_ENABLED_FOR_EXTERNAL_CONTRIBUTORS_ONLY \
  --build-config=cloudbuild-pr.yaml \
  --service-account=projects/acme/serviceAccounts/cb-checkout-pr@acme.iam.gserviceaccount.com

--included-files and --ignored-files are the monorepo essentials. Without them every commit builds every service; with them a README change builds nothing. The patterns are matched against the commit’s changed files, and getting them slightly wrong (missing a shared library path) produces the opposite problem — a service that does not rebuild when its dependency changed. Test both directions.

--comment-control is the fork trust boundary: COMMENTS_ENABLED_FOR_EXTERNAL_CONTRIBUTORS_ONLY means a maintainer must comment /gcbrun before a fork’s PR builds. For a public repository this is not optional — a fork PR can change cloudbuild.yaml itself.

Trigger approvals add a second gate, independent of the build config:

gcloud builds triggers create github --name=deploy-prod --require-approval …
gcloud builds approve BUILD_ID --project=acme

Useful when the trigger itself is the privileged operation, and audited in Cloud Audit Logs.

Other trigger types worth knowing exist: Pub/Sub (chain a build off another build, or off an Artifact Registry push), webhook (anything external), and manual — including gcloud builds triggers run for a rebuild, which is how you re-run a release without pushing an empty commit.


Topic 2: Identity — One Service Account Per Trigger

This is the single most impactful operational decision in Cloud Build.

The legacy default service account accumulates permissions across every pipeline in the project. A build triggered by anything — including a pull request — inherits all of them. One service account per trigger, scoped to what that pipeline does:

TriggerNeeds
PR buildlogging.logWriter, read from Artifact Registry. No write, no deploy.
main build+ artifactregistry.writer, secretmanager.secretAccessor on its own secrets
release build+ clouddeploy.releaser, iam.serviceAccountUser on the deploy SA
anythingcloudbuild.builds.builder on itself
gcloud projects add-iam-policy-binding acme \
  --member=serviceAccount:cb-checkout-pr@acme.iam.gserviceaccount.com \
  --role=roles/logging.logWriter

A required setting that surprises people: when you specify a user service account, the build needs somewhere to write logs. Either grant roles/logging.logWriter and set options: logging: CLOUD_LOGGING_ONLY, or provide a logs bucket. A build that fails immediately with a logging permission error is almost always this.

Cross-project pushes — building in one project and pushing to a shared registry in another — are a grant on the registry project, not a key:

gcloud artifacts repositories add-iam-policy-binding apps \
  --location=europe-west1 --project=acme-artifacts \
  --member=serviceAccount:cb-checkout-main@acme.iam.gserviceaccount.com \
  --role=roles/artifactregistry.writer

No stored credential anywhere in that flow, which is the structural advantage over a self-hosted CI holding cloud keys.


Topic 3: Worker Pools and Networking

"THE BUILD CANNOT REACH IT" IS ALMOST ALWAYS THE POOL DEFAULT POOL Google-managed public internet your VPC ✗ No route to private resources. No fixed egress IP. Free tier, fastest to start, fine for public dependencies. PRIVATE POOL peered network your VPC ✓ Cloud NAT → fixed IP Reaches private databases, internal registries, on-prem. Billed per pool; choose the machine type deliberately. MACHINE TYPE IS A SPEED DECISION, NOT ONLY A COST ONE E2_MEDIUM (default) · E2_HIGHCPU_8 · E2_HIGHCPU_32 · with diskSizeGb for large checkouts Billing is per build-minute, so a build that halves its duration on a bigger machine can cost the same and return feedback twice as fast. Measure both before choosing.
The default pool has no route into your VPC — which is the cause of most 'the build times out reaching our database' reports. The bottom panel is the tuning decision people skip.
gcloud builds worker-pools create private-pool \
  --region=europe-west1 \
  --peered-network=projects/acme/global/networks/prod-vpc \
  --peered-network-ip-range=/24 \
  --worker-machine-type=e2-standard-8 \
  --worker-disk-size=200 \
  --no-public-egress
options:
  pool:
    name: 'projects/acme/locations/europe-west1/workerPools/private-pool'

What each flag decides:

  • --peered-network gives builds a route into your VPC — internal databases, private registries, on-premises over interconnect.
  • --no-public-egress removes direct internet access. Pair it with Cloud NAT for a fixed egress IP, which is what a partner’s allow-list needs.
  • --worker-disk-size matters for large checkouts and big image builds; the default fills faster than expected on a monorepo.

Regional placement matters for both latency and data residency: put the pool where your registry and clusters are, or every image push crosses a region.

Machine type is a speed decision. Billing is per build-minute, so a build that takes 8 minutes on E2_MEDIUM and 3 on E2_HIGHCPU_8 may cost roughly the same and returns feedback nearly three times faster. Measure both before choosing — this is one of the few CI tuning knobs where the fast option is often not the expensive one.


Topic 4: Notifications and Observability

Builds publish to a Pub/Sub topic called cloud-builds automatically. Everything else is built on that:

gcloud pubsub subscriptions create build-failures \
  --topic=cloud-builds \
  --message-filter='attributes.status="FAILURE"'

The official notifier images (Slack, SMTP, HTTP, BigQuery) consume that topic and are deployed as Cloud Run services with a small config:

apiVersion: cloud-build-notifiers/v1
kind: SlackNotifier
metadata: { name: slack-failures }
spec:
  notification:
    filter: build.status == Build.Status.FAILURE
    delivery:
      webhookUrl:
        secretRef: slack-webhook

Logs go to Cloud Logging by default, with the same retention, export and alerting as everything else in the project:

gcloud builds log BUILD_ID
gcloud builds log BUILD_ID --stream          # follow a running build

# Every failed build in the last day, with the trigger that started it
gcloud builds list --filter='status=FAILURE AND createTime>-P1D' \
  --format='table(id, substitutions.TRIGGER_NAME, duration)'

Export builds to BigQuery if you want DORA metrics from real data: a Pub/Sub subscription writing to BigQuery gives you deployment frequency and lead time as queries rather than estimates. The first lesson asked for those four numbers; this is the least-effort way to get two of them.


Topic 5: Cost, Quotas and the Limits That Bite

Cost is per build-minute by machine type, plus the private pool’s own hourly rate if you use one, plus logs and artifact storage. The line item that surprises teams is not the compute — it is a monorepo trigger with no includedFiles, building twelve services on every commit.

Quotas worth knowing before you meet them:

concurrent builds        per project, per region (adjustable)
build timeout            default 10 min, maximum 24 h — set it explicitly
max build steps          100
max artifacts            100 per build
worker pool size         per pool, and it queues rather than failing

The default 10-minute timeout is the most common first surprise. A build that dies at exactly 600 seconds with no useful error is the timeout, and timeout: 1800s at the top level plus per-step timeouts is the fix.

Concurrency queues rather than failing, which means a monorepo with aggressive triggers produces a queue that looks like slowness. gcloud builds list --ongoing tells you whether you are compute-bound or queue-bound — different problems, different fixes.


Topic 6: Managing It as Code

Triggers configured by hand in the console have the same problem as freestyle Jenkins jobs: no review, no diff, no history.

resource "google_cloudbuild_trigger" "checkout_main" {
  name            = "checkout-main"
  location        = "europe-west1"
  service_account = google_service_account.cb_checkout_main.id
  filename        = "cloudbuild.yaml"

  included_files = ["services/checkout/**"]
  ignored_files  = ["**/*.md"]

  github {
    owner = "acme"
    name  = "checkout"
    push { branch = "^main$" }
  }

  substitutions = { _ENVIRONMENT = "staging" }
}

Two more practices that keep a Cloud Build estate reviewable:

  • Keep the build config in the repository, not inline in the trigger. Inline configs are invisible to the developers whose code they build — the same argument as the Jenkinsfile.
  • Validate before pushing. There is no server-side linter, but gcloud builds submit --no-source --config=cloudbuild.yaml fails fast on a malformed config, and a YAML schema check in a pre-commit hook catches the rest.

When Cloud Build is the wrong tool, stated plainly so the next lesson makes sense: it is a DAG of steps with no conditionals, no matrix, no human input step and no loop. Multi-environment promotion, approvals and canaries do not belong in cloudbuild.yaml — and every team that tries ends up writing a worse Cloud Deploy in bash. That is the next two lessons.

Try it yourself: give a PR trigger a service account with no artifactregistry.writer and watch the push step fail with a clear IAM error. That failure is your fork trust boundary working, and it is far better to see it deliberately than to discover it was never there.

Common mistake: leaving every trigger on the default Cloud Build service account because it “just works”. It accumulates permissions from every pipeline in the project, so a compromised dependency in any build — including an unreviewed fork’s pull request — inherits the union of them. One service account per trigger takes ten minutes and bounds the blast radius permanently.