IP Addressing and the Gateway API

Ephemeral versus reserved, regional versus global, why an apex record forces you to reserve, and giving a GKE Gateway a stable address.

intermediate 20 min lesson hands-on task included

IP addresses are the least interesting part of a design until a recreate changes one and every DNS record, firewall allow-list and partner integration points at nothing.


Topic 1: Two Axes That Decide Everything

TWO AXES: EPHEMERAL vs RESERVED, AND REGIONAL vs GLOBAL EPHEMERAL RESERVED (static) lifetime released when the resource goes yours until you delete it survives recreate no — new address yes — same address DNS A record unusable this is the one to use billed when idle n/a yes — unattached costs money GLOBAL vs REGIONAL — NOT INTERCHANGEABLE --global → global external ALB, only --region= → VMs, regional LBs, Cloud NAT, PSC endpoints THE GATEWAY API TAKES ONE TOO addresses: [{ type: NamedAddress, value: web-ip }] Reserve first, name it in the Gateway — or the IP changes. UNATTACHED STATIC IPs ARE BILLED, AND THEY ACCUMULATE gcloud compute addresses list --filter='status=RESERVED' --format='table(name,region,address,status)' ← run this monthly
Ephemeral versus reserved decides whether the address survives; regional versus global decides what can use it. The bottom banner is the monthly clean-up nobody schedules.
# Global — only the global external load balancer uses these
gcloud compute addresses create web-ip --global

# Regional — VMs, regional load balancers, Cloud NAT, PSC endpoints
gcloud compute addresses create nat-ip-1 --region=europe-west1

# Internal — a fixed private address inside a subnet
gcloud compute addresses create db-endpoint \
  --region=europe-west1 --subnet=data-subnet --purpose=GCE_ENDPOINT

Ephemeral addresses are allocated when a resource is created and released when it goes. Fine for a VM nobody connects to by address; useless for anything referenced elsewhere.

Reserved (static) addresses are yours until you delete them. They survive recreating the resource, which is what makes them usable in a DNS record, a partner allow-list or a firewall rule on someone else’s network.

Global and regional are not interchangeable, and the error when you mix them is unhelpfully generic. A global external ALB needs --global; everything else needs --region=.


Topic 2: Why the Apex Record Forces a Reservation

From the DNS lesson: you cannot CNAME a zone apex. So acme.example must be an A record pointing at a literal address — which means that address must never change.

reserve a global static IP
      ↓
create the load balancer forwarding rule against it
      ↓
create the apex A record pointing at it
      ↓
recreate the load balancer freely; the address survives

That chain is the practical reason most production load balancers sit on reserved addresses, and it is worth building in that order rather than discovering it after a rebuild changed the IP.

The same logic applies to egress. A partner that allow-lists your outbound address needs Cloud NAT with reserved addresses, not automatic allocation:

gcloud compute routers nats create prod-nat --router=nat-router \
  --region=europe-west1 --nat-external-ip-pool=nat-ip-1,nat-ip-2 \
  --nat-all-subnet-ip-ranges

Topic 3: The Gateway API

The Gateway API is the successor to Ingress and the direction GKE networking is going. It splits one overloaded object into three with different owners:

# Platform team owns this
apiVersion: gateway.networking.k8s.io/v1
kind: Gateway
metadata:
  name: external-gateway
  namespace: infra
spec:
  gatewayClassName: gke-l7-global-external-managed
  addresses:
    - type: NamedAddress
      value: web-ip                 # the reserved global address, by name
  listeners:
    - name: https
      protocol: HTTPS
      port: 443
      tls:
        mode: Terminate
        options:
          networking.gke.io/pre-shared-certs: acme-cert
      allowedRoutes:
        namespaces:
          from: Selector
          selector:
            matchLabels: { gateway-access: "true" }
# Application team owns this, in their own namespace
apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata:
  name: checkout
  namespace: payments
spec:
  parentRefs: [{ name: external-gateway, namespace: infra }]
  hostnames: ["checkout.acme.example"]
  rules:
    - matches: [{ path: { type: PathPrefix, value: /api } }]
      backendRefs:
        - name: checkout
          port: 8080
          weight: 90
        - name: checkout-canary
          port: 8080
          weight: 10

Three properties that make this better than Ingress:

  • Role separation. The platform team owns the Gateway, its address and its certificates; application teams attach routes without touching either. allowedRoutes is the boundary.
  • Real weighted traffic splitting — the weight field above is what a genuine canary needs, and what the CI/CD module’s progressive delivery lesson assumes.
  • Portable. The API is upstream Kubernetes; the gatewayClassName is the GCP-specific part.

The addresses field is the reason this lesson pairs them. Without it the Gateway takes an ephemeral IP and a recreate changes it. With a NamedAddress pointing at a reserved address, the Gateway can be deleted and recreated and DNS never notices.

Gateway classes worth knowing:

gke-l7-global-external-managed    global anycast, multi-cluster capable
gke-l7-regional-external-managed  regional, cheaper
gke-l7-rilb                       internal, for east-west traffic

Topic 4: Certificates on a Gateway

Two mechanisms, and the newer one is worth adopting:

# Pre-shared: a Compute Engine SSL certificate, referenced by name
options:
  networking.gke.io/pre-shared-certs: acme-cert
# Certificate Manager: better for many domains and wildcards
apiVersion: networking.gke.io/v1
kind: ManagedCertificate
metadata: { name: acme-cert }
spec:
  domains: [checkout.acme.example]

Both require the DNS record to already point at the Gateway’s address before provisioning completes — which is the third reason the address gets reserved first. A certificate stuck in PROVISIONING is almost always a DNS record that does not yet resolve to the right IP.


Topic 5: Internal Addressing

Not every address is external. Inside the VPC:

  • Primary subnet range — what VMs get.
  • Secondary ranges (alias IPs) — what GKE pods and services get, which is why the GKE lesson insisted on sizing them before cluster creation.
  • Reserved internal addresses — a fixed private IP for a database endpoint, an internal load balancer, or a PSC endpoint, so that internal DNS records stay valid.
gcloud compute addresses create ilb-ip \
  --region=europe-west1 --subnet=app-subnet \
  --addresses=10.20.0.100 --purpose=SHARED_LOADBALANCER_VIP

Reserving an internal address you have chosen (rather than letting GCP pick) is worth doing for anything referenced in configuration, because it makes the value predictable across environments.


Topic 6: The Monthly Audit

Reserved addresses are billed when unattached — GCP charges for holding an address you are not using, deliberately, to discourage hoarding.

# Every reserved address, and whether anything uses it
gcloud compute addresses list \
  --format='table(name, region.basename(), address, status, users.len())'

# The ones costing money for nothing
gcloud compute addresses list --filter='status=RESERVED' \
  --format='table(name, region.basename(), address, creationTimestamp)'

status=IN_USE is attached; status=RESERVED is idle and billed. A project that has been through a few migrations usually has several, and they are pure waste.

And check for the opposite problem: a DNS record pointing at an address that is no longer reserved by you. That is the subdomain-takeover risk from the DNS lesson, and this list is how you find it.

Try it yourself: create a Gateway without an addresses field, note the IP, delete and recreate it, and note the new one. Then add a reserved NamedAddress and repeat. The second IP does not change — which is the difference between a DNS record you can trust and one you cannot.

Common mistake: letting a load balancer or Gateway take an ephemeral address in a staging environment “because it does not matter”, then copying that pattern into production. It matters the first time the resource is recreated — a Terraform replacement, a cluster rebuild, a region migration — and the outage lasts as long as the old TTL, plus however long it takes someone to work out why.