Four GCP features get described as “making things private”, and they solve genuinely different problems. Choosing the wrong one produces a design that looks secure and is not, or a network that cannot reach anything for reasons nobody can find.
Topic 1: Four Mechanisms, Four Problems
| Mechanism | Answers | Operates on |
|---|---|---|
| Private Google Access | How does a VM with no external IP call a Google API? | A subnet flag |
| Cloud NAT | How does it reach the wider internet? | Managed egress translation |
| Private Service Connect | How do I consume a service over a private address? | An endpoint IP in your VPC |
| VPC Service Controls | How do I stop data leaving, even by an authorised identity? | A perimeter around APIs |
They compose rather than compete: a hardened private subnet typically has all four.
Topic 2: Private Google Access and Cloud NAT
Private Google Access is one flag on a subnet. With it, instances with no external IP reach Google APIs (storage.googleapis.com, logging.googleapis.com) through Google’s internal network:
gcloud compute networks subnets update app-subnet \
--region=europe-west1 --enable-private-ip-google-access
Without it, a private VM cannot write a log line, pull an image from Artifact Registry, or read a secret — and the symptom is a hang followed by a timeout, not a clear error. It is free, and there is no reason for a private subnet not to have it.
private.googleapis.com vs restricted.googleapis.com is the detail that matters when a perimeter is involved: the restricted VIP (199.36.153.4/30) only resolves services that VPC Service Controls supports, which prevents an exfiltration path to a non-supported API. Inside a perimeter, route Google API traffic there:
# Cloud DNS private zone for googleapis.com
*.googleapis.com → CNAME restricted.googleapis.com
restricted.googleapis.com → A 199.36.153.4, .5, .6, .7
Cloud NAT covers everything that is not a Google API — package repositories, third-party APIs, webhooks:
gcloud compute routers create nat-router --network=prod-vpc --region=europe-west1
gcloud compute routers nats create prod-nat \
--router=nat-router --region=europe-west1 \
--nat-all-subnet-ip-ranges --nat-external-ip-pool=nat-ip-1,nat-ip-2 \
--enable-logging --min-ports-per-vm=128
Three properties to plan around:
- It is egress only. There is no inbound path, which is the point.
- Reserve static IPs rather than using automatic allocation, because a partner allow-list needs an address that does not change.
- Port exhaustion is the failure mode. Many VMs talking to one destination exhaust the port pool; the symptom is intermittent connection failures that correlate with load.
min-ports-per-vmand dynamic port allocation are the levers, and NAT logging is how you see it coming.
Topic 3: Private Service Connect
PSC gives you an endpoint with an IP address in your own subnet that forwards to a service — Google’s, a partner’s, or another team’s.
# Consume a published service privately
gcloud compute addresses create psc-sql --region=europe-west1 --subnet=app-subnet
gcloud compute forwarding-rules create psc-sql-endpoint \
--region=europe-west1 --network=prod-vpc --address=psc-sql \
--target-service-attachment=projects/p/regions/europe-west1/serviceAttachments/sql-attachment
Why this beats VPC peering for consuming a service:
- No CIDR coordination. The provider’s address space is irrelevant, and overlap does not matter.
- One direction. The consumer reaches the service; the service cannot reach into the consumer’s VPC.
- No transitive routing surprises, and no route table changes as consumers are added.
- It works across organisations, which peering makes awkward.
The mirror image — publishing a service — is a service attachment in front of an internal load balancer, with an explicit accept-list of consumer projects. This is how a platform team offers an internal service to twenty product teams without twenty peerings.
Private Service Access (the older VPC-peering-based mechanism, still used by some managed services) allocates a range you reserve for Google to peer into. Where a service supports both, PSC is the direction of travel.
Topic 4: VPC Service Controls
Everything above is about network reachability. VPC-SC is about data movement by an authorised identity — the case where credentials are valid and the request should still be refused.
gcloud access-context-manager perimeters create prod-perimeter \
--title="Production data perimeter" \
--resources=projects/111111111111,projects/222222222222 \
--restricted-services=storage.googleapis.com,bigquery.googleapis.com \
--policy=POLICY_ID
What it stops that IAM does not: a principal with legitimate storage.objectViewer copying an object from a bucket inside the perimeter to a project outside it. IAM says yes; the perimeter says no, because the destination is outside.
The parts that make or break a rollout:
- Dry-run first, for weeks.
--dry-runlogs every request that would be denied, without denying it. VPC-SC blocks in ways nothing else in GCP does, and the violations are always more numerous than expected.
gcloud logging read 'protoPayload.metadata.dryRun=true AND
protoPayload.metadata."@type"="type.googleapis.com/google.cloud.audit.VpcServiceControlAuditMetadata"' \
--limit=50 --format='table(protoPayload.authenticationInfo.principalEmail, protoPayload.methodName, resource.labels.project_id)'
- Ingress and egress rules are how legitimate crossings are allowed — a specific service account, from a specific project, to a specific service. Keep them narrow and commented.
- Access levels add conditions on the caller: corporate IP range, device policy, identity. This is where BeyondCorp-style access lands.
- Your own CI breaks first. A pipeline outside the perimeter that reads a bucket inside it is the most common dry-run finding, and the fix is an ingress rule for that service account rather than widening the perimeter.
Topic 5: Hybrid Connectivity, and Cloud DNS
HA VPN — two tunnels, IPsec over the internet, ~99.99% with both. Fast to set up, bandwidth bounded by tunnel and internet path.
Dedicated Interconnect — a physical circuit into a Google colocation facility. 10 or 100 Gbps, predictable latency, weeks to provision. Partner Interconnect goes through a service provider when you cannot reach a facility.
Neither is encrypted by default — Interconnect is a private circuit, not an encrypted one. A compliance requirement for encryption in transit needs HA VPN over the Interconnect, or MACsec where available. That sentence catches people out on every cloud, and GCP is no exception.
Cloud DNS ties the private estate together:
# A private zone, resolvable only inside the VPC
gcloud dns managed-zones create internal \
--dns-name=internal.acme.example. --visibility=private --networks=prod-vpc
# Forward a corporate zone to on-premises resolvers
gcloud dns managed-zones create corp-forward \
--dns-name=corp.acme.example. --visibility=private --networks=prod-vpc \
--forwarding-targets=10.99.0.10,10.99.0.11
DNS peering lets a service project resolve names published in the host project’s private zone — the piece that makes Shared VPC feel like one network rather than several.
Topic 6: A Private Subnet That Actually Works
The end state, and the order to build it in:
□ subnet with --enable-private-ip-google-access
□ no external IPs (enforced by org policy, not by convention)
□ Cloud NAT with reserved static IPs and logging on
□ PSC endpoints for the managed services you consume
□ firewall: default deny egress, allow by network tag, with a reason per rule
□ private Cloud DNS zone, plus forwarding to on-prem where needed
□ VPC-SC perimeter in dry-run, then enforced, with narrow ingress rules
□ VPC flow logs on, sampled, exported where you can query them
Then verify with Connectivity Tests rather than by trying it, which analyses the configuration path and names the blocking component:
gcloud network-management connectivity-tests create app-to-sql \
--source-instance=projects/p/zones/europe-west1-b/instances/app-1 \
--destination-ip-address=10.20.130.5 --destination-port=5432 --protocol=TCP
gcloud network-management connectivity-tests describe app-to-sql \
--format='value(reachabilityDetails.result, reachabilityDetails.traces)'
It costs nothing, needs no packets, and replaces a great deal of guessing about which of firewall, route or perimeter is responsible.
Try it yourself: build a VM with no external IP and turn the four mechanisms on one at a time, testing after each. The sequence — nothing works, Google APIs work, the internet works, the database works privately — makes each mechanism’s job unmistakable.
Common mistake: enforcing a VPC Service Controls perimeter without a dry-run period, on the assumption that it only blocks external attackers. It blocks your CI pipeline, your BigQuery scheduled queries, your Terraform runner and the analytics team’s notebooks — all authorised, all now outside the perimeter. Two weeks of dry-run logs turns that from an outage into a list of ingress rules.