Cloud Load Balancing, CDN and Cloud Armor

One anycast IP for the planet, the five objects between it and your backends, and the health-check firewall rule that causes most 502s.

intermediate 24 min lesson hands-on task included

GCP’s global load balancer is the most distinctive thing in its networking stack: one anycast IP advertised from every Google edge, with no DNS failover and no per-region address. Understanding the objects behind it is what makes it debuggable.


Topic 1: The Five Objects

ONE ANYCAST IP, EVERY REGION — THE PIECES BETWEEN IT AND YOUR PODS anycast IP one address, advertised globally forwarding rule IP + port → target proxy target proxy terminates TLS (ACM-style certs) URL map host and path → backend service backend service health checks, CDN, Armor, affinity backends (NEGs) MIGs, GKE pods, Cloud Run, on-prem WHAT GLOBAL BUYS YOU · one IP for the whole planet — no DNS failover · users enter at the nearest Google edge · cross-region overflow when a backend is full · Cloud CDN and Cloud Armor attach here Regional LBs exist, are cheaper — and are regional. WHERE IT GOES WRONG · health check firewall rule missing → all backends unhealthy, 502 everywhere · 35.191.0.0/16 and 130.211.0.0/22 must be allowed to reach your instances · cert provisioning takes minutes, not seconds · propagation is global; changes are not instant THE FIRST THING TO CHECK ON ANY 502 gcloud compute backend-services get-health BACKEND --global — all UNHEALTHY = the health check, or its firewall rule
Read the chain top to bottom — a request touches all five. The panel on the right is the failure mode that accounts for most first-time 502s, and the command at the bottom is how you confirm it in one line.
anycast IP  →  forwarding rule  →  target proxy  →  URL map  →  backend service  →  backends

Forwarding rule — binds the IP and port to a target proxy. This is the object that holds the address.

Target proxy — terminates the connection and TLS. Google-managed certificates attach here.

URL map — routes by host and path to a backend service. This is where /api and /static diverge.

Backend service — the important one. It owns the health check, session affinity, timeouts, Cloud CDN, Cloud Armor, and the balancing mode.

Backends — instance groups or network endpoint groups (NEGs). A NEG can be GKE pods (GCE_VM_IP_PORT), Cloud Run (SERVERLESS), or an on-premises endpoint (INTERNET_FQDN_PORT) — which is how one load balancer fronts a hybrid estate.

gcloud compute backend-services create checkout-be \
  --global --protocol=HTTP --port-name=http \
  --health-checks=checkout-hc --enable-cdn \
  --connection-draining-timeout=60

Topic 2: Global vs Regional, and Which to Choose

Global external ALBRegional external ALBInternal ALBNetwork LB (L4)
IPOne anycast, worldwideOne per regionInternal, per regionRegional, pass-through
Layer7774
Cross-region failoverAutomaticNoNoNo
Cloud CDNYesNoNoNo
Cloud ArmorYesYesLimitedNo
Preserves client IPVia X-Forwarded-ForVia headerVia headerYes, natively

Global is the default choice for anything public, and the reason is failover: when a region’s backends go unhealthy, the same IP serves the next-nearest region with no DNS change and no TTL to wait out. That single property removes an entire class of DR machinery.

Reach for the L4 network load balancer when you need the real client IP without a header, a non-HTTP protocol, or the lowest possible latency. Note the trade — with a pass-through LB, your firewall rules must allow the client ranges, not the load balancer’s.

Balancing mode decides when a backend is considered full and traffic spills to the next region:

RATE          requests per second per instance or endpoint
UTILIZATION   backend CPU (instance groups only)
CONNECTION    concurrent connections (L4)

RATE with a maxRatePerEndpoint you have actually measured is what makes overflow behave predictably. Left at defaults, a region absorbs far more than it should before spilling.


Topic 3: Health Checks — and the Firewall Rule

This is the single most common GCP load-balancing failure, and it is worth stating on its own: Google’s health-check probes come from 35.191.0.0/16 and 130.211.0.0/22. If a firewall rule does not allow those ranges to your backend port, every backend reports UNHEALTHY, the load balancer has nowhere to send traffic, and every request returns 502.

gcloud compute firewall-rules create allow-health-checks \
  --network=prod-vpc --direction=INGRESS --action=allow \
  --source-ranges=35.191.0.0/16,130.211.0.0/22 \
  --target-tags=http-backend --rules=tcp:8080
# The first command to run on any 502
gcloud compute backend-services get-health checkout-be --global

Health check design, which matters as much as the firewall rule:

gcloud compute health-checks create http checkout-hc \
  --port=8080 --request-path=/healthz \
  --check-interval=10s --timeout=5s \
  --healthy-threshold=2 --unhealthy-threshold=3
  • The path must be cheap and must not check dependencies. A /healthz that queries the database turns one database blip into every backend going unhealthy at once — the same rule as Kubernetes readiness probes.
  • timeout must be less than check-interval.
  • Unhealthy threshold × interval is how long a genuinely dead backend keeps receiving traffic. Three checks at ten seconds is thirty seconds of errors.

Topic 4: TLS and Cloud CDN

Google-managed certificates are free, auto-renewing, and the right default:

gcloud compute ssl-certificates create checkout-cert \
  --domains=checkout.acme.example,www.acme.example --global

Two operational notes: provisioning requires the DNS A record to already point at the load balancer IP and takes up to about an hour, and a certificate covering a domain that stops resolving will eventually fail to renew. Both surface as PROVISIONING that never becomes ACTIVE.

Cloud CDN is a flag on the backend service, not a separate product:

gcloud compute backend-services update checkout-be --global \
  --enable-cdn --cache-mode=CACHE_ALL_STATIC \
  --default-ttl=3600 --max-ttl=86400 --negative-caching
Cache modeBehaviour
USE_ORIGIN_HEADERSRespects your Cache-Control — correct, and requires your app to be right
CACHE_ALL_STATICCaches static content types automatically
FORCE_CACHE_ALLCaches everything, including responses that should never be cached

FORCE_CACHE_ALL on a path that serves per-user content is how one user’s session lands in another user’s browser. Use USE_ORIGIN_HEADERS and fix the application’s headers, or scope FORCE_CACHE_ALL to a genuinely static path prefix.

Cache invalidation is slow and rate-limited — design around versioned URLs (/static/app.9f8e7d.js) rather than relying on purging.


Topic 5: Cloud Armor

A security policy attaches to a backend service and evaluates rules in priority order before traffic reaches your backends.

gcloud compute security-policies create checkout-armor \
  --description="edge protection for checkout"

# Managed protection against the OWASP top ten
gcloud compute security-policies rules create 1000 \
  --security-policy=checkout-armor \
  --expression="evaluatePreconfiguredExpr('xss-v33-stable')" --action=deny-403

# Rate limiting per client IP
gcloud compute security-policies rules create 2000 \
  --security-policy=checkout-armor \
  --src-ip-ranges='*' --action=rate-based-ban \
  --rate-limit-threshold-count=100 --rate-limit-threshold-interval-sec=60 \
  --ban-duration-sec=600 --conform-action=allow --exceed-action=deny-429 \
  --enforce-on-key=IP

# Geographic restriction
gcloud compute security-policies rules create 3000 \
  --security-policy=checkout-armor \
  --expression="origin.region_code == 'XX'" --action=deny-403

Deploy every rule in preview mode first:

gcloud compute security-policies rules update 1000 \
  --security-policy=checkout-armor --preview

Preview logs what the rule would have blocked without blocking it. The preconfigured WAF rules have real false-positive rates against ordinary application traffic — a rule that blocks a legitimate JSON body looks exactly like an application bug to everyone except the person reading the Armor logs.

Adaptive Protection (Managed Protection Plus) learns baseline traffic and proposes rules during an attack. Worth knowing it exists; worth enabling if you are a plausible DDoS target.


Topic 6: Debugging the Path

# 1. Are the backends healthy? Almost always the answer.
gcloud compute backend-services get-health checkout-be --global

# 2. What did the load balancer actually log?
gcloud logging read \
  'resource.type="http_load_balancer" AND httpRequest.status>=500' \
  --limit=20 --format='table(timestamp, httpRequest.status, jsonPayload.statusDetails)'

# 3. Is the certificate live?
gcloud compute ssl-certificates describe checkout-cert --global \
  --format='value(managed.status, managed.domainStatus)'

# 4. Is the path open at all, per the configuration?
gcloud network-management connectivity-tests create lb-to-backend \
  --source-ip-address=35.191.0.1 --destination-instance=… --protocol=TCP --destination-port=8080

jsonPayload.statusDetails is the field that names the cause, and knowing four values covers most incidents:

statusDetailsMeaning
failed_to_pick_backendNo healthy backend — health check or firewall
backend_connection_closed_before_data_sent_to_clientYour app closed the connection; often a keep-alive shorter than the LB’s
response_sent_by_backendThe 5xx came from your application, not the LB
denied_by_security_policyCloud Armor blocked it

That table turns “the load balancer is returning errors” into a specific component in one query, which is the difference between a five-minute diagnosis and an afternoon.

Try it yourself: delete the health-check firewall rule on a working setup and watch every backend flip to UNHEALTHY and every request return 502 within about thirty seconds. Restoring the rule fixes it just as fast. Doing this once makes the failure instantly recognisable.

Common mistake: setting the backend service timeout to the default 30 seconds for an endpoint that legitimately takes longer — a report, an upload, a streaming response. The load balancer cuts the connection mid-response and the client sees a truncated body or a 502, while the application log shows the request completing successfully. The two halves never appear in the same place, which is why this one takes so long to find.