The build config says what runs. This lesson is everything around it: what starts a build, what identity it carries, what network it sits in, and what it costs — which is where the operational decisions actually are.
Topic 1: Triggers
# Push to a branch
gcloud builds triggers create github \
--name=checkout-main \
--repo-owner=acme --repo-name=checkout \
--branch-pattern='^main$' \
--build-config=cloudbuild.yaml \
--included-files='services/checkout/**' \
--ignored-files='**/*.md,docs/**' \
--service-account=projects/acme/serviceAccounts/cb-checkout-main@acme.iam.gserviceaccount.com \
--substitutions=_ENVIRONMENT=staging
# Tag push — the production path
gcloud builds triggers create github \
--name=checkout-release \
--tag-pattern='^v[0-9]+\.[0-9]+\.[0-9]+$' \
--build-config=cloudbuild-release.yaml \
--service-account=projects/acme/serviceAccounts/cb-checkout-release@acme.iam.gserviceaccount.com
# Pull request — build and test only
gcloud builds triggers create github \
--name=checkout-pr \
--pull-request-pattern='^main$' \
--comment-control=COMMENTS_ENABLED_FOR_EXTERNAL_CONTRIBUTORS_ONLY \
--build-config=cloudbuild-pr.yaml \
--service-account=projects/acme/serviceAccounts/cb-checkout-pr@acme.iam.gserviceaccount.com
--included-files and --ignored-files are the monorepo essentials. Without them every commit builds every service; with them a README change builds nothing. The patterns are matched against the commit’s changed files, and getting them slightly wrong (missing a shared library path) produces the opposite problem — a service that does not rebuild when its dependency changed. Test both directions.
--comment-control is the fork trust boundary: COMMENTS_ENABLED_FOR_EXTERNAL_CONTRIBUTORS_ONLY means a maintainer must comment /gcbrun before a fork’s PR builds. For a public repository this is not optional — a fork PR can change cloudbuild.yaml itself.
Trigger approvals add a second gate, independent of the build config:
gcloud builds triggers create github --name=deploy-prod --require-approval …
gcloud builds approve BUILD_ID --project=acme
Useful when the trigger itself is the privileged operation, and audited in Cloud Audit Logs.
Other trigger types worth knowing exist: Pub/Sub (chain a build off another build, or off an Artifact Registry push), webhook (anything external), and manual — including gcloud builds triggers run for a rebuild, which is how you re-run a release without pushing an empty commit.
Topic 2: Identity — One Service Account Per Trigger
This is the single most impactful operational decision in Cloud Build.
The legacy default service account accumulates permissions across every pipeline in the project. A build triggered by anything — including a pull request — inherits all of them. One service account per trigger, scoped to what that pipeline does:
| Trigger | Needs |
|---|---|
| PR build | logging.logWriter, read from Artifact Registry. No write, no deploy. |
| main build | + artifactregistry.writer, secretmanager.secretAccessor on its own secrets |
| release build | + clouddeploy.releaser, iam.serviceAccountUser on the deploy SA |
| anything | cloudbuild.builds.builder on itself |
gcloud projects add-iam-policy-binding acme \
--member=serviceAccount:cb-checkout-pr@acme.iam.gserviceaccount.com \
--role=roles/logging.logWriter
A required setting that surprises people: when you specify a user service account, the build needs somewhere to write logs. Either grant roles/logging.logWriter and set options: logging: CLOUD_LOGGING_ONLY, or provide a logs bucket. A build that fails immediately with a logging permission error is almost always this.
Cross-project pushes — building in one project and pushing to a shared registry in another — are a grant on the registry project, not a key:
gcloud artifacts repositories add-iam-policy-binding apps \
--location=europe-west1 --project=acme-artifacts \
--member=serviceAccount:cb-checkout-main@acme.iam.gserviceaccount.com \
--role=roles/artifactregistry.writer
No stored credential anywhere in that flow, which is the structural advantage over a self-hosted CI holding cloud keys.
Topic 3: Worker Pools and Networking
gcloud builds worker-pools create private-pool \
--region=europe-west1 \
--peered-network=projects/acme/global/networks/prod-vpc \
--peered-network-ip-range=/24 \
--worker-machine-type=e2-standard-8 \
--worker-disk-size=200 \
--no-public-egress
options:
pool:
name: 'projects/acme/locations/europe-west1/workerPools/private-pool'
What each flag decides:
--peered-networkgives builds a route into your VPC — internal databases, private registries, on-premises over interconnect.--no-public-egressremoves direct internet access. Pair it with Cloud NAT for a fixed egress IP, which is what a partner’s allow-list needs.--worker-disk-sizematters for large checkouts and big image builds; the default fills faster than expected on a monorepo.
Regional placement matters for both latency and data residency: put the pool where your registry and clusters are, or every image push crosses a region.
Machine type is a speed decision. Billing is per build-minute, so a build that takes 8 minutes on E2_MEDIUM and 3 on E2_HIGHCPU_8 may cost roughly the same and returns feedback nearly three times faster. Measure both before choosing — this is one of the few CI tuning knobs where the fast option is often not the expensive one.
Topic 4: Notifications and Observability
Builds publish to a Pub/Sub topic called cloud-builds automatically. Everything else is built on that:
gcloud pubsub subscriptions create build-failures \
--topic=cloud-builds \
--message-filter='attributes.status="FAILURE"'
The official notifier images (Slack, SMTP, HTTP, BigQuery) consume that topic and are deployed as Cloud Run services with a small config:
apiVersion: cloud-build-notifiers/v1
kind: SlackNotifier
metadata: { name: slack-failures }
spec:
notification:
filter: build.status == Build.Status.FAILURE
delivery:
webhookUrl:
secretRef: slack-webhook
Logs go to Cloud Logging by default, with the same retention, export and alerting as everything else in the project:
gcloud builds log BUILD_ID
gcloud builds log BUILD_ID --stream # follow a running build
# Every failed build in the last day, with the trigger that started it
gcloud builds list --filter='status=FAILURE AND createTime>-P1D' \
--format='table(id, substitutions.TRIGGER_NAME, duration)'
Export builds to BigQuery if you want DORA metrics from real data: a Pub/Sub subscription writing to BigQuery gives you deployment frequency and lead time as queries rather than estimates. The first lesson asked for those four numbers; this is the least-effort way to get two of them.
Topic 5: Cost, Quotas and the Limits That Bite
Cost is per build-minute by machine type, plus the private pool’s own hourly rate if you use one, plus logs and artifact storage. The line item that surprises teams is not the compute — it is a monorepo trigger with no includedFiles, building twelve services on every commit.
Quotas worth knowing before you meet them:
concurrent builds per project, per region (adjustable)
build timeout default 10 min, maximum 24 h — set it explicitly
max build steps 100
max artifacts 100 per build
worker pool size per pool, and it queues rather than failing
The default 10-minute timeout is the most common first surprise. A build that dies at exactly 600 seconds with no useful error is the timeout, and timeout: 1800s at the top level plus per-step timeouts is the fix.
Concurrency queues rather than failing, which means a monorepo with aggressive triggers produces a queue that looks like slowness. gcloud builds list --ongoing tells you whether you are compute-bound or queue-bound — different problems, different fixes.
Topic 6: Managing It as Code
Triggers configured by hand in the console have the same problem as freestyle Jenkins jobs: no review, no diff, no history.
resource "google_cloudbuild_trigger" "checkout_main" {
name = "checkout-main"
location = "europe-west1"
service_account = google_service_account.cb_checkout_main.id
filename = "cloudbuild.yaml"
included_files = ["services/checkout/**"]
ignored_files = ["**/*.md"]
github {
owner = "acme"
name = "checkout"
push { branch = "^main$" }
}
substitutions = { _ENVIRONMENT = "staging" }
}
Two more practices that keep a Cloud Build estate reviewable:
- Keep the build config in the repository, not inline in the trigger. Inline configs are invisible to the developers whose code they build — the same argument as the Jenkinsfile.
- Validate before pushing. There is no server-side linter, but
gcloud builds submit --no-source --config=cloudbuild.yamlfails fast on a malformed config, and a YAML schema check in a pre-commit hook catches the rest.
When Cloud Build is the wrong tool, stated plainly so the next lesson makes sense: it is a DAG of steps with no conditionals, no matrix, no human input step and no loop. Multi-environment promotion, approvals and canaries do not belong in cloudbuild.yaml — and every team that tries ends up writing a worse Cloud Deploy in bash. That is the next two lessons.
Try it yourself: give a PR trigger a service account with no artifactregistry.writer and watch the push step fail with a clear IAM error. That failure is your fork trust boundary working, and it is far better to see it deliberately than to discover it was never there.
Common mistake: leaving every trigger on the default Cloud Build service account because it “just works”. It accumulates permissions from every pipeline in the project, so a compromised dependency in any build — including an unreviewed fork’s pull request — inherits the union of them. One service account per trigger takes ten minutes and bounds the blast radius permanently.