The GCP Failure Playbook

A six-step sweep for any GCP incident, telling an IAM denial from an org policy denial, and the quota and propagation failures that present as random.

advanced 22 min lesson hands-on task included

Under pressure people debug by intuition, and intuition goes to the thing they understand best. A fixed sequence beats it because it eliminates whole categories in a known order — and on GCP two of those categories are unique enough to be worth naming.


Topic 1: The Sweep

FIRST SIX QUESTIONS — ANSWER ALL OF THEM BEFORE FORMING A THEORY 1 Who am I, in which project? gcloud config list gcloud auth list the one people skip 2 Is it Google or is it us? Service Health dashboard Personalized Service Health 3 What changed in the last hour? Cloud Audit Logs protoPayload.methodName 4 Is it IAM or an org policy? Policy Troubleshooter org-policies describe --effective 5 Are we at a quota? gcloud compute project-info describe Quotas page, per region 6 Is the path open? Connectivity Tests firewall-rules list + VPC flow logs WHY STEP 4 IS SPECIFIC TO GCP A denial can come from IAM or from an org policy, and the messages look similar. The Policy Troubleshooter tells you which — guessing wastes the most time here.
Answer all six before forming a theory. Step 4 is the GCP-specific one — a denial can come from IAM or from an org policy, the messages look alike, and guessing wastes the most time.
# 1. Who am I, in which project?
gcloud config list
gcloud auth list

# 2. Is it Google, or is it us?
#    console.cloud.google.com/servicehealth — and Personalized Service Health
gcloud logging read 'resource.type="gce_instance" severity>=ERROR' --limit=10 --freshness=1h

# 3. What changed in the last hour?
gcloud logging read \
  'protoPayload."@type"="type.googleapis.com/google.cloud.audit.AuditLog"
   AND protoPayload.methodName!~"^google.*(list|get)"' \
  --limit=30 --freshness=1h \
  --format='table(timestamp, protoPayload.authenticationInfo.principalEmail, protoPayload.methodName)'

# 4. Is it IAM, or an org policy?
gcloud policy-troubleshoot iam PROJECT_RESOURCE \
  --principal-email=svc@acme.iam.gserviceaccount.com --permission=storage.objects.get
gcloud resource-manager org-policies describe CONSTRAINT --project=PROJECT --effective

# 5. Are we at a quota?
gcloud compute project-info describe --format='value(quotas)'
#    plus the Quotas page, which is per service AND per region

# 6. Is the path open?
gcloud network-management connectivity-tests create t1 --source-instance=… --destination-ip-address=… --protocol=TCP --destination-port=443

Steps 2 and 3 are the highest-yield pair: most incidents are either Google having a problem or somebody having changed something, and both are answerable in a minute.

Record what normal looks like. A sweep whose healthy output you have never seen is much harder to read under pressure.


Topic 2: IAM Denial or Org Policy Denial

This is the GCP-specific fork, and the error messages are genuinely similar.

PERMISSION_DENIED: Required 'compute.instances.create' permission
   → IAM. A role is missing.

Constraint constraints/compute.vmExternalIpAccess violated for project X
   → ORG POLICY. No role fixes this; the policy must change or an exception exists.

The Policy Troubleshooter answers the IAM half definitively — it evaluates every binding, including inherited ones and conditions, and tells you which granted or failed to grant:

gcloud policy-troubleshoot iam \
  //cloudresourcemanager.googleapis.com/projects/checkout-prod \
  --principal-email=deployer@acme.iam.gserviceaccount.com \
  --permission=compute.instances.create

For the org policy half, --effective is the only reliable read, because the answer is assembled from the organisation, folders and project:

gcloud resource-manager org-policies describe \
  compute.vmExternalIpAccess --project=checkout-prod --effective

Three more denial sources worth recognising, because they present the same way:

  • VPC Service Controls — the error mentions a perimeter or securityPolicyViolation. Check the VPC-SC audit logs; the fix is an ingress rule, not a role.
  • A service not enabled — SERVICE_DISABLED. gcloud services enable ….
  • Domain-restricted sharing — granting a role to an external identity fails because of iam.allowedPolicyMemberDomains.

Topic 3: Quotas

Almost every GCP resource has a limit, most are per project per region, and the ones that bite are rarely the ones people check.

gcloud compute regions describe europe-west1 --format='table(quotas.metric, quotas.usage, quotas.limit)'
gcloud services quota list --service=compute.googleapis.com --consumer=projects/acme

The usual suspects:

CPUS / CPUS_ALL_REGIONS        stops a scale-out cold
IN_USE_ADDRESSES               external IPs, per region
SSD_TOTAL_GB                   persistent disk
Cloud SQL instances per project
Cloud Run: concurrent requests / instances per region
GKE: nodes per cluster, pods per node
API rate quotas — per minute, per user, and easy to hit from a loop

Request increases before you need them. An increase can take hours or days and always arrives during the event you needed it for. Alarm on quota utilisation at 80% — Cloud Monitoring exposes serviceruntime.googleapis.com/quota/allocation/usage for exactly this.

Quotas are per project, which is one more argument for the project-per-workload split from the hierarchy lesson: a load test that exhausts a quota in its own project does not stop production scaling.


Topic 4: The Failures That Present as Random

Eventual consistency in the control plane. An IAM binding takes seconds to propagate — sometimes longer. Automation that grants a role and immediately uses it fails intermittently, and the retry succeeds. This is why Terraform occasionally needs a second apply.

Firewall rule priority. Rules are evaluated by priority number, lowest first, and the implied deny all ingress at 65535 is easy to forget. A rule that “should work” is often shadowed by a lower-numbered deny.

gcloud compute firewall-rules list --sort-by=priority \
  --format='table(name, priority, direction, sourceRanges.list(), allowed[].map().firewall_rule().list())'

Health-check source ranges. 35.191.0.0/16 and 130.211.0.0/22 must be allowed, or every backend is unhealthy and everything returns 502 — covered in the load balancing lesson and worth repeating because it is the single most common GCP networking incident.

Cloud NAT port exhaustion. Many VMs to one destination exhausts the port pool; connections fail intermittently under load. nat_allocation_failed in the NAT metrics is the tell.

Private Google Access not enabled. A private VM hangs on every Google API call, then times out. No error names the cause.

Organisation policy applied to existing resources. The resources keep running and become unmanageable — you cannot modify or recreate them, and they keep billing.


Topic 5: Instance and Workload Triage

When a VM is unreachable, in order of what needs the least to work:

# 1. Is it running, and what did the guest say on boot?
gcloud compute instances describe vm-1 --zone=europe-west1-b --format='value(status)'
gcloud compute instances get-serial-port-output vm-1 --zone=europe-west1-b | tail -40

# 2. In without SSH — IAP TCP forwarding needs no external IP and no bastion
gcloud compute ssh vm-1 --zone=europe-west1-b --tunnel-through-iap

# 3. Run a diagnostic without a shell
gcloud compute instances add-metadata vm-1 --zone=… --metadata-from-file startup-script=diag.sh

# 4. Last resort: detach the boot disk and attach it to a working VM

IAP TCP forwarding is the GCP equivalent of Session Manager and deserves to be the default: no external IP, no bastion, no open port 22, and every session is IAM-authorised and audited. If your VMs still have public IPs for SSH access, this is the change that removes them.

For GKE workloads the triage is ordinary Kubernetes triage — the kubectl triage cheat sheet covers the sweep, and the GCP-specific additions are node pool status and the cluster’s own operations:

gcloud container operations list --filter='status=RUNNING'

Topic 6: Logging Queries Worth Keeping

# Every non-read admin action, by principal, in the last hour
gcloud logging read \
  'logName:"cloudaudit.googleapis.com%2Factivity"' --freshness=1h \
  --format='table(timestamp, protoPayload.authenticationInfo.principalEmail, protoPayload.methodName, resource.labels.project_id)'

# Denials, whatever their source
gcloud logging read 'protoPayload.status.code=7' --freshness=1h \
  --format='table(timestamp, protoPayload.authenticationInfo.principalEmail, protoPayload.methodName, protoPayload.status.message)'

# Load balancer 5xx with the reason
gcloud logging read 'resource.type="http_load_balancer" AND httpRequest.status>=500' \
  --freshness=1h --format='table(timestamp, httpRequest.status, jsonPayload.statusDetails)'

# What a specific service account has been doing
gcloud logging read \
  'protoPayload.authenticationInfo.principalEmail="svc@acme.iam.gserviceaccount.com"' \
  --freshness=24h --limit=50

protoPayload.status.code=7 is PERMISSION_DENIED and it is the single most useful filter in an access incident — it surfaces IAM, org policy and VPC-SC denials together, with the principal and the method that failed.

Save these as saved queries or a log-based dashboard before you need them; writing a filter correctly under pressure is harder than it looks.

Try it yourself: run all six sweep steps against a healthy project and save the output. Two minutes now turns the first two minutes of your next incident into pattern matching rather than exploration.

Common mistake: treating a PERMISSION_DENIED as an IAM problem and adding roles until it works. On GCP the denial may be an org policy, a VPC Service Controls perimeter, a disabled service or domain-restricted sharing — none of which any role fixes. You end up with an over-privileged service account and the original problem, because nobody read which of the four actually refused.