Cloud DNS: Zones, Records and Routing

Public, private, forwarding and peering zones, the record types you will actually write, why you cannot CNAME an apex, and TTL as your rollback speed.

intermediate 20 min lesson hands-on task included

DNS is where a working system becomes reachable, and it is the layer where a five-minute change takes an hour to take effect. Both halves matter operationally.


Topic 1: Four Zone Types

FOUR ZONE TYPES — THE VISIBILITY DECIDES WHO CAN RESOLVE IT Public the internet resolves it delegate with NS records at the registrar Private only the attached VPCs internal names, split-horizon safe Forwarding ask someone else on-prem resolvers for a corp zone Peering use another VPC’s zones service project resolves host names RECORD TYPES YOU WILL ACTUALLY WRITE A / AAAA → an address CNAME → another name MX → mail TXT → verification, SPF, DKIM NS → delegation CAA → who may issue certificates THE APEX CNAME RULE You cannot CNAME the zone apex (example.com). Use an A record to a reserved static IP — which is why load balancer IPs get reserved, not ephemeral. TTL IS YOUR ROLLBACK SPEED A 3600s TTL means an hour before a DNS change fully takes effect. Lower it to 60s DAYS BEFORE a planned migration, not during it. And remember some resolvers ignore short TTLs — which is why a global anycast IP beats DNS failover whenever it is available.
Visibility decides who can resolve a zone. The bottom banner is the operational one — TTL is the speed at which you can undo a mistake.
# Public — the internet resolves it
gcloud dns managed-zones create acme-public \
  --dns-name=acme.example. --visibility=public --description="public zone"

# Private — only the attached VPCs
gcloud dns managed-zones create internal \
  --dns-name=internal.acme.example. --visibility=private --networks=prod-vpc

# Forwarding — send these queries to another resolver
gcloud dns managed-zones create corp \
  --dns-name=corp.acme.example. --visibility=private --networks=prod-vpc \
  --forwarding-targets=10.99.0.10,10.99.0.11

# Peering — resolve using another VPC's zones
gcloud dns managed-zones create host-peer \
  --dns-name=internal.acme.example. --visibility=private --networks=svc-vpc \
  --target-network=prod-vpc --target-project=host-project

Peering zones are what make Shared VPC feel like one network. A service project’s workloads can resolve names published in the host project’s private zone without republishing anything.

Split horizon is the pattern where the same name resolves differently inside and outside: a public zone returning the load balancer IP, and a private zone for the same name returning an internal address. Useful, and worth documenting loudly — a name that resolves to two things is confusing at 3am.


Topic 2: Records You Will Actually Write

gcloud dns record-sets create app.acme.example. \
  --zone=acme-public --type=A --ttl=300 --rrdatas=34.120.0.10

gcloud dns record-sets create www.acme.example. \
  --zone=acme-public --type=CNAME --ttl=300 --rrdatas=app.acme.example.

gcloud dns record-sets create acme.example. \
  --zone=acme-public --type=MX --ttl=3600 \
  --rrdatas="10 mail1.acme.example.,20 mail2.acme.example."
TypeForNote
A / AAAAAn addressThe workhorse
CNAMEAn alias to another nameNever at the zone apex
MXMail routingPriority then host
TXTVerification, SPF, DKIM, DMARCQuoting matters
NSDelegating a subdomainAlso how the zone itself is delegated
SRVService discoveryPort and weight
CAAWhich CAs may issue certificates for youCheap, and few people set it

The trailing dot is not optional. app.acme.example. is fully qualified; without the dot, some tooling appends the zone name and you get app.acme.example.acme.example.

The apex CNAME rule: DNS does not permit a CNAME at a zone apex alongside the SOA and NS records that must be there. So acme.example cannot be a CNAME to a load balancer hostname — it must be an A record to an address, which is exactly why load balancer IPs get reserved rather than left ephemeral.


Topic 3: Routing Policies

Cloud DNS can answer differently per client, which covers cases the global load balancer does not:

# Geolocation — answer by where the query came from
gcloud dns record-sets create api.acme.example. --zone=acme-public --type=A --ttl=60 \
  --routing-policy-type=GEO \
  --routing-policy-data="europe-west1=34.1.1.1;us-east1=35.2.2.2"

# Weighted round robin — a DNS-level canary
gcloud dns record-sets create api.acme.example. --zone=acme-public --type=A --ttl=60 \
  --routing-policy-type=WRR \
  --routing-policy-data="90.0=34.1.1.1;10.0=35.2.2.2"

# Failover with health checking
gcloud dns record-sets create api.acme.example. --zone=acme-public --type=A --ttl=30 \
  --routing-policy-type=FAILOVER \
  --routing-policy-primary-data=34.1.1.1 --routing-policy-backup-data=35.2.2.2 \
  --enable-health-checking

Prefer the global load balancer where you can. A single anycast IP fails over in a health-check interval with no client caching involved; DNS failover waits for TTLs that some resolvers ignore. DNS routing earns its place for non-HTTP protocols, for multi-cloud, and for splitting traffic to endpoints a single LB cannot front.


Topic 4: TTL Is Your Rollback Speed

A TTL is a promise to resolvers that they may cache the answer for that long. It is therefore the floor on how quickly a change — including a fix — takes effect.

TTL 3600   an hour before everyone sees the change
TTL 300    five minutes; a reasonable default
TTL 60     a minute; use during a migration

Lower the TTL days before a planned change, not during it. Dropping from 3600 to 60 takes up to an hour to be picked up, because resolvers are still holding the old TTL. Lower it, wait, migrate, then raise it again.

Some resolvers ignore short TTLs, and browsers cache independently. Plan for a long tail of clients still using the old answer, which is another reason a stable IP beats a DNS change whenever both are options.


Topic 5: DNSSEC and Zone Hygiene

gcloud dns managed-zones update acme-public --dnssec-state=on
gcloud dns dns-keys list --zone=acme-public   # then give the DS record to your registrar

DNSSEC signs your responses so a resolver can detect tampering. It is one flag plus a DS record at the registrar, and the failure mode is worth respecting: a broken DNSSEC chain makes the domain fail to resolve entirely, which is a harder outage than a wrong record. Enable it deliberately and verify the DS record before considering it done.

Zone hygiene, which prevents the two most common DNS incidents:

  • Dangling records. An A record pointing at an IP you released can be claimed by someone else — subdomain takeover. Audit records against live resources.
  • Records nobody owns. Manage zones in Terraform so every record has a commit and an author.
# Every record in a zone, for review
gcloud dns record-sets list --zone=acme-public --format='table(name, type, ttl, rrdatas.list())'

Topic 6: Debugging Resolution

# From inside the VPC — the metadata resolver at 169.254.169.254
dig app.internal.acme.example @169.254.169.254 +short
dig +trace acme.example        # follow delegation from the root

# What Cloud DNS actually holds, bypassing every cache
gcloud dns record-sets list --zone=internal --name=app.internal.acme.example.

The order to check when a name does not resolve:

1. Does the record exist in the zone?           record-sets list
2. Is the zone attached to this VPC?            managed-zones describe
3. Is the query even reaching Cloud DNS?        dig against 169.254.169.254
4. Is a forwarding or peering zone shadowing it? more specific zone wins
5. Is it cached locally?                        check the TTL, and the client's own cache

Step 4 is the GCP-specific one. Zone selection is longest-suffix-match, so a forwarding zone for acme.example. will capture app.internal.acme.example. unless a more specific private zone exists. Two zones covering overlapping suffixes is a configuration to review carefully.

Try it yourself: create a private zone, resolve a name from a VM inside the VPC, then try the same query from your laptop. One works and one does not — which is the whole point of private zones and takes thirty seconds to demonstrate.

Common mistake: pointing a production A record at an ephemeral IP. The address is released when the resource is recreated, the record now points at nothing — or worse, at an address someone else has since been allocated. Reserve the address first, then create the record, and audit for records whose target no longer exists.