Serverless: Cloud Run and Cloud Functions

Concurrency as the thing that makes Cloud Run cheap, cold starts and what actually fixes them, revisions and traffic splitting, and when a cluster is still the answer.

intermediate 24 min lesson hands-on task included

Cloud Run is the default place to run a container on GCP, and the two settings that decide whether it is cheap and fast are ones most teams never change.


Topic 1: Where Code Runs on GCP

FOUR PLACES TO RUN CODE — ORDERED BY HOW MUCH YOU OPERATE CLOUD FUNCTIONS CLOUD RUN GKE AUTOPILOT GKE / GCE unit a function a container a pod a node scales to zero yes yes no no concurrency 1 per instance up to 1000 you decide you decide cold start yes yes (min-instances helps) no no you operate nothing nothing workloads everything best at event glue HTTP services + jobs many services anything unusual THE DEFAULT ANSWER ON GCP IS CLOUD RUN It takes a container, scales to zero, splits traffic between revisions, and has no cluster. Reach for GKE when you need what a cluster gives you.
Ordered by how much you operate. The banner at the bottom is the honest default — reach past Cloud Run only when you need what a cluster gives you.

Cloud Run takes a container, gives it an HTTPS endpoint, scales it from zero, and charges per request-second. No cluster, no nodes, no capacity planning.

Cloud Functions (2nd gen) is Cloud Run underneath with a source-to-container build step and event triggers attached. If you are choosing today, choose Cloud Run and add an Eventarc trigger — you get the same eventing with a portable container.

GKE Autopilot is the answer when you need pod-level primitives: DaemonSets, sidecars with lifecycle guarantees, StatefulSets, a service mesh, or many services sharing one network policy model.

Compute Engine is for anything unusual: a licence tied to a host, a GPU workload with specific drivers, a process that must run for days.

gcloud run deploy checkout \
  --image=europe-west1-docker.pkg.dev/acme/apps/checkout@sha256:9f8e7d... \
  --region=europe-west1 \
  --service-account=checkout@acme.iam.gserviceaccount.com \
  --no-allow-unauthenticated \
  --concurrency=80 --cpu=1 --memory=512Mi \
  --min-instances=1 --max-instances=100

Topic 2: Concurrency Is the Setting That Matters

This is the difference between Cloud Run and every function-as-a-service that preceded it: one instance can handle many requests at once.

concurrency = 1     one request per instance. 100 concurrent requests = 100 instances.
concurrency = 80    one instance serves 80. 100 concurrent requests ≈ 2 instances.

Since billing is per instance-second, that is close to a 40× cost difference for the same traffic — and it is why a Cloud Run bill that looks like Lambda’s usually means concurrency was left at 1.

What concurrency requires from your application: it must be genuinely concurrent. A single-threaded, blocking process with concurrency=80 accepts 80 requests and serves them one at a time, and the 80th waits. Node, Go and async Python are fine; a synchronous WSGI app with one worker is not — give it a real worker count first.

How to pick a number: start at 80, load-test, and watch p99 latency and memory. Raise it until latency degrades or memory pressure appears, then back off. Set concurrency=1 only when the workload genuinely cannot share — a process that holds a global lock, or heavy per-request memory.

CPU allocation is the paired decision:

  • CPU during request processing only (default) — cheapest, and background work after the response stops abruptly.
  • CPU always allocated — required for background threads, connection pool maintenance and streaming after response. Costs more per instance-second.

A common bug: a service that writes telemetry after returning the response, deployed with the default, silently loses those writes.


Topic 3: Cold Starts, and What Actually Fixes Them

A cold start is container pull plus process start plus first-request warm-up. What genuinely helps, in order of effect:

--min-instances — keep N instances warm. This is the fix; everything else is optimisation. One warm instance costs roughly a small VM and removes cold starts from the p99 for normal traffic.

A smaller image. Pull time is real. A 1 GB image with a full JDK starts far more slowly than a 90 MB distroless one — the container module’s multi-stage build lesson is directly relevant.

--cpu-boost — extra CPU during startup, which measurably helps JVM and .NET services.

Lazy initialisation done deliberately. Work at import time is paid on every cold start; work on first use is paid once per instance. Connection pools and clients should be created at startup and reused, not per request.

gcloud run services update checkout --min-instances=2 --cpu-boost

Measure before tuning. Cloud Run exports run.googleapis.com/container/startup_latencies; a p95 of 300ms needs no work, and a p95 of 8 seconds is usually image size plus framework startup, not the platform.


Topic 4: Revisions and Traffic Splitting

Every deploy creates an immutable revision. Traffic is assigned to revisions explicitly, which gives you canary and instant rollback without any extra machinery:

# Deploy without taking traffic
gcloud run deploy checkout --image=…@sha256:… --no-traffic --tag=candidate

# Test the candidate directly on its own URL
curl https://candidate---checkout-abc123-ew.a.run.app/healthz

# 10% canary
gcloud run services update-traffic checkout --to-tags=candidate=10

# Promote, or roll back instantly
gcloud run services update-traffic checkout --to-latest
gcloud run services update-traffic checkout --to-revisions=checkout-00042-abc=100

Rollback is a traffic change, so it takes seconds and needs no rebuild — the old revision is still there. That property makes the “measure rollback, not deployment” rule from the CI/CD module easy to satisfy here.

Deploy by digest, not by tag (@sha256:…). A revision pinned to a moving tag is not the immutable artifact it appears to be.


Topic 5: Networking, Identity and Data

Ingress decides who can reach the service at the network level:

--ingress=all                        # public
--ingress=internal                   # VPC and VPC-SC only
--ingress=internal-and-cloud-load-balancing   # behind your global LB — the common production choice

Authentication is separate from ingress. --no-allow-unauthenticated requires an IAM identity with run.invoker; combine both for a service that is neither publicly routable nor publicly callable.

Egress into your VPC — for a private Cloud SQL, an internal API, or an on-premises system — uses Direct VPC egress or a Serverless VPC Access connector:

gcloud run deploy checkout \
  --network=prod-vpc --subnet=serverless-subnet \
  --vpc-egress=private-ranges-only

all-traffic routes internet egress through your VPC too, which is how a Cloud Run service gets a fixed egress IP via Cloud NAT — usually because a partner allow-lists addresses.

Cloud SQL connects either through that private path or with the built-in connector (--add-cloudsql-instances), which uses IAM rather than an authorised-network list.

Secrets mount from Secret Manager as environment variables or files (--set-secrets), so no credential is baked into the image or the service definition.


Topic 6: Jobs, Eventarc, and the Limits

Cloud Run jobs run to completion rather than serving requests — migrations, batch, scheduled work:

gcloud run jobs create migrate --image=…@sha256:… --tasks=1 --max-retries=3
gcloud run jobs execute migrate --wait
gcloud scheduler jobs create http nightly-report --schedule="0 2 * * *" …

Eventarc delivers events from ~100 Google sources — a GCS object finalised, a Pub/Sub message, an Audit Log entry — to a Cloud Run service as an HTTP request, which is what replaces most 1st-gen Cloud Functions triggers.

The limits worth knowing before you design around Cloud Run:

request timeout      up to 60 minutes (default 5) — but a load balancer in front has its own
instance memory      up to 32 GiB
CPU                  up to 8 vCPU
concurrency          up to 1000 per instance
no persistent local disk — /tmp is in-memory and counts against your memory limit
no fixed IP without VPC egress + Cloud NAT

When to leave Cloud Run for GKE, honestly: you need DaemonSets or sidecars, you are running many services that share network policy and mesh configuration, you need long-lived stateful workloads, or the per-request pricing has crossed over well-packed nodes at steady high load. Below that, a cluster is operational cost you are paying for nothing.

Try it yourself: deploy the same service with --concurrency=1 and --concurrency=80, then run identical load against both and compare instance count in the metrics. The gap is the single largest cost lever on Cloud Run, and seeing it once changes how you deploy.

Common mistake: treating Cloud Run instances as long-lived and keeping state in memory — a cache, a session, an in-progress upload. Instances are created and destroyed constantly, and with --min-instances=0 the last one disappears entirely between requests. Anything that must survive belongs in Memorystore, a database or GCS, and discovering that in production usually looks like an intermittent bug rather than a design error.