Cloud Run is the default place to run a container on GCP, and the two settings that decide whether it is cheap and fast are ones most teams never change.
Topic 1: Where Code Runs on GCP
Cloud Run takes a container, gives it an HTTPS endpoint, scales it from zero, and charges per request-second. No cluster, no nodes, no capacity planning.
Cloud Functions (2nd gen) is Cloud Run underneath with a source-to-container build step and event triggers attached. If you are choosing today, choose Cloud Run and add an Eventarc trigger — you get the same eventing with a portable container.
GKE Autopilot is the answer when you need pod-level primitives: DaemonSets, sidecars with lifecycle guarantees, StatefulSets, a service mesh, or many services sharing one network policy model.
Compute Engine is for anything unusual: a licence tied to a host, a GPU workload with specific drivers, a process that must run for days.
gcloud run deploy checkout \
--image=europe-west1-docker.pkg.dev/acme/apps/checkout@sha256:9f8e7d... \
--region=europe-west1 \
--service-account=checkout@acme.iam.gserviceaccount.com \
--no-allow-unauthenticated \
--concurrency=80 --cpu=1 --memory=512Mi \
--min-instances=1 --max-instances=100
Topic 2: Concurrency Is the Setting That Matters
This is the difference between Cloud Run and every function-as-a-service that preceded it: one instance can handle many requests at once.
concurrency = 1 one request per instance. 100 concurrent requests = 100 instances.
concurrency = 80 one instance serves 80. 100 concurrent requests ≈ 2 instances.
Since billing is per instance-second, that is close to a 40× cost difference for the same traffic — and it is why a Cloud Run bill that looks like Lambda’s usually means concurrency was left at 1.
What concurrency requires from your application: it must be genuinely concurrent. A single-threaded, blocking process with concurrency=80 accepts 80 requests and serves them one at a time, and the 80th waits. Node, Go and async Python are fine; a synchronous WSGI app with one worker is not — give it a real worker count first.
How to pick a number: start at 80, load-test, and watch p99 latency and memory. Raise it until latency degrades or memory pressure appears, then back off. Set concurrency=1 only when the workload genuinely cannot share — a process that holds a global lock, or heavy per-request memory.
CPU allocation is the paired decision:
- CPU during request processing only (default) — cheapest, and background work after the response stops abruptly.
- CPU always allocated — required for background threads, connection pool maintenance and streaming after response. Costs more per instance-second.
A common bug: a service that writes telemetry after returning the response, deployed with the default, silently loses those writes.
Topic 3: Cold Starts, and What Actually Fixes Them
A cold start is container pull plus process start plus first-request warm-up. What genuinely helps, in order of effect:
--min-instances — keep N instances warm. This is the fix; everything else is optimisation. One warm instance costs roughly a small VM and removes cold starts from the p99 for normal traffic.
A smaller image. Pull time is real. A 1 GB image with a full JDK starts far more slowly than a 90 MB distroless one — the container module’s multi-stage build lesson is directly relevant.
--cpu-boost — extra CPU during startup, which measurably helps JVM and .NET services.
Lazy initialisation done deliberately. Work at import time is paid on every cold start; work on first use is paid once per instance. Connection pools and clients should be created at startup and reused, not per request.
gcloud run services update checkout --min-instances=2 --cpu-boost
Measure before tuning. Cloud Run exports run.googleapis.com/container/startup_latencies; a p95 of 300ms needs no work, and a p95 of 8 seconds is usually image size plus framework startup, not the platform.
Topic 4: Revisions and Traffic Splitting
Every deploy creates an immutable revision. Traffic is assigned to revisions explicitly, which gives you canary and instant rollback without any extra machinery:
# Deploy without taking traffic
gcloud run deploy checkout --image=…@sha256:… --no-traffic --tag=candidate
# Test the candidate directly on its own URL
curl https://candidate---checkout-abc123-ew.a.run.app/healthz
# 10% canary
gcloud run services update-traffic checkout --to-tags=candidate=10
# Promote, or roll back instantly
gcloud run services update-traffic checkout --to-latest
gcloud run services update-traffic checkout --to-revisions=checkout-00042-abc=100
Rollback is a traffic change, so it takes seconds and needs no rebuild — the old revision is still there. That property makes the “measure rollback, not deployment” rule from the CI/CD module easy to satisfy here.
Deploy by digest, not by tag (@sha256:…). A revision pinned to a moving tag is not the immutable artifact it appears to be.
Topic 5: Networking, Identity and Data
Ingress decides who can reach the service at the network level:
--ingress=all # public
--ingress=internal # VPC and VPC-SC only
--ingress=internal-and-cloud-load-balancing # behind your global LB — the common production choice
Authentication is separate from ingress. --no-allow-unauthenticated requires an IAM identity with run.invoker; combine both for a service that is neither publicly routable nor publicly callable.
Egress into your VPC — for a private Cloud SQL, an internal API, or an on-premises system — uses Direct VPC egress or a Serverless VPC Access connector:
gcloud run deploy checkout \
--network=prod-vpc --subnet=serverless-subnet \
--vpc-egress=private-ranges-only
all-traffic routes internet egress through your VPC too, which is how a Cloud Run service gets a fixed egress IP via Cloud NAT — usually because a partner allow-lists addresses.
Cloud SQL connects either through that private path or with the built-in connector (--add-cloudsql-instances), which uses IAM rather than an authorised-network list.
Secrets mount from Secret Manager as environment variables or files (--set-secrets), so no credential is baked into the image or the service definition.
Topic 6: Jobs, Eventarc, and the Limits
Cloud Run jobs run to completion rather than serving requests — migrations, batch, scheduled work:
gcloud run jobs create migrate --image=…@sha256:… --tasks=1 --max-retries=3
gcloud run jobs execute migrate --wait
gcloud scheduler jobs create http nightly-report --schedule="0 2 * * *" …
Eventarc delivers events from ~100 Google sources — a GCS object finalised, a Pub/Sub message, an Audit Log entry — to a Cloud Run service as an HTTP request, which is what replaces most 1st-gen Cloud Functions triggers.
The limits worth knowing before you design around Cloud Run:
request timeout up to 60 minutes (default 5) — but a load balancer in front has its own
instance memory up to 32 GiB
CPU up to 8 vCPU
concurrency up to 1000 per instance
no persistent local disk — /tmp is in-memory and counts against your memory limit
no fixed IP without VPC egress + Cloud NAT
When to leave Cloud Run for GKE, honestly: you need DaemonSets or sidecars, you are running many services that share network policy and mesh configuration, you need long-lived stateful workloads, or the per-request pricing has crossed over well-packed nodes at steady high load. Below that, a cluster is operational cost you are paying for nothing.
Try it yourself: deploy the same service with --concurrency=1 and --concurrency=80, then run identical load against both and compare instance count in the metrics. The gap is the single largest cost lever on Cloud Run, and seeing it once changes how you deploy.
Common mistake: treating Cloud Run instances as long-lived and keeping state in memory — a cache, a session, an in-progress upload. Instances are created and destroyed constantly, and with --min-instances=0 the last one disappears entirely between requests. Anything that must survive belongs in Memorystore, a database or GCS, and discovering that in production usually looks like an intermittent bug rather than a design error.