GCP Operations
Master GCP operations: Resource Hierarchy & Labels, IAM & Service Account delegation, Shared VPCs & Firewall Tags, GCE MIGs & SUD/CUD cost optimization, GCS WORM Lifecycles, GKE Pod IP & Workload Identity, Cloud Logging Sinks, and Pub/Sub & BigQuery analytics.
Stage 1 β Hierarchy & Security
5 lessonsHow GCP structures resources through Organizations, Folders, and Projects, and why understanding the difference between Labels and Network Tags prevents severe security and billing mistakes.
Master GCP Cloud IAM bindings, Primitive vs Predefined roles, Service Account security, and eliminating exported JSON keys with Service Account Impersonation.
Constraints that cap what any project may do, why an org policy grants nothing, and how to roll one out in dry-run before it blocks a production deploy.
The three levels of key control, what CMEK actually changes about a managed service, and why the key becomes a dependency you have to operate.
Versions as immutable payloads, why disabling comes before destroying, replication and residency, rotation notifications, and consuming a secret with no key file anywhere.
Stage 2 β Networking & Shared Topologies
5 lessonsUnderstand GCP's unique global VPC network design, regional subnets, firewall rule targeting via Network Tags, and enterprise Shared VPC topologies.
One anycast IP for the planet, the five objects between it and your backends, and the health-check firewall rule that causes most 502s.
Private Google Access, Cloud NAT, Private Service Connect and VPC Service Controls β four mechanisms that solve four different problems and are constantly confused.
Public, private, forwarding and peering zones, the record types you will actually write, why you cannot CNAME an apex, and TTL as your rollback speed.
Ephemeral versus reserved, regional versus global, why an apex record forces you to reserve, and giving a GKE Gateway a stable address.
Stage 3 β Compute & Storage Operations
7 lessonsMaster GCE machine families, Custom machine types, Managed Instance Groups (MIGs), auto-healing, Spot VMs, and automated vs committed cost optimization.
Master GCS object storage classes, automated lifecycle transitions, bucket retention locks for WORM compliance, and Uniform Bucket-Level Access.
Concurrency as the thing that makes Cloud Run cheap, cold starts and what actually fixes them, revisions and traffic splitting, and when a cluster is still the answer.
HA versus read replicas versus backups β three features that solve three different problems β plus the Auth Proxy, maintenance windows, and choosing between Cloud SQL, AlloyDB, Spanner and Firestore.
Images, disks and snapshots, what metadata actually controls, OS Login and IAP instead of SSH keys, and what survives a stop, a delete and a live migration.
Compute separated from distributed storage, read pools that scale independently, the columnar engine, and an honest account of when Cloud SQL is still the right answer.
Turning an object change into a message, why every consumer must be idempotent, the loop that costs money, and choosing between Pub/Sub notifications and Eventarc.
Stage 4 β Kubernetes, Observability & Analytics
8 lessonsMaster Google Kubernetes Engine (GKE) Autopilot vs Standard, VPC-Native Alias IP networking, Node Pool management, and pod identity delegation via Workload Identity.
Master GCP Cloud Logging, Audit vs Data Access logs, building Log Router Sinks to BigQuery & GCS, log-based metrics, and Cloud Monitoring alert policies.
Master Cloud Pub/Sub vs Cloud Tasks messaging, BigQuery partitioning/clustering optimization, managed databases, VPC Service Controls, and zero-trust IAP access.
Upgrades happen whether you plan them or not β release channels, maintenance windows, surge behaviour, and the Autopilot-versus-Standard cost question answered with numbers.
Where a log entry physically lives, log buckets and the _Default/_Required split, analytics views, the query language worth learning properly, and how retention becomes a bill.
How routing actually evaluates, the four destinations and what each is for, aggregated sinks at the organisation level, and the writer identity that makes exports silently fail.
Push versus pull, ack deadlines and the redelivery loop, ordering keys and what they cost, dead-letter topics, exactly-once, and the one metric that tells you a consumer is broken.
Reading the database's replication log instead of querying it, the backfill-then-stream model, connectivity choices, schema drift, and when CDC is the wrong answer.
Stage 5 β Reliability & Capstone
3 lessonsZonal, regional and multi-regional resource scope, turning availability into an RTO and an RPO, and why a global load balancer makes regional failover a health-check outcome.
A six-step sweep for any GCP incident, telling an IAM denial from an org policy denial, and the quota and propagation failures that present as random.
Build the whole module as one environment β org policies, Shared VPC, private workloads, CMEK, a global load balancer β then break it seven ways and prove it recovers.
πΊοΈ Beginner β Expert Roadmap
5 stages with prerequisites and a concrete mastery check at each.
π― What You'll Learn
- β’ Organize GCP infrastructure with Organizations, Folders, Projects, and differentiate Labels from Network Tags.
- β’ Implement Cloud IAM least privilege, predefined vs primitive roles, and Service Account Impersonation without long-lived keys.
- β’ Design Global VPC networks, Regional Subnets, Firewall rules with Network Tags, and Shared VPC topologies.
- β’ Optimize GCE machine families, Custom machine types, Spot VMs, Sustained Use Discounts (SUD), and Committed Use Discounts (CUD).
- β’ Configure GCS storage classes (Standard/Nearline/Coldline/Archive), Bucket Lifecycle policies, and immutable Bucket Locks.
- β’ Deploy GKE clusters with VPC-Native Alias IP networking, custom Node Pools, and pod-level IAM via Workload Identity.
- β’ Build Cloud Logging Sinks to GCS, BigQuery, and Pub/Sub, configure log-based metrics, and set up Cloud Monitoring alerts.
- β’ Architect event-driven pipelines with Pub/Sub vs Cloud Tasks, BigQuery partitioning/clustering, and zero-trust IAP/VPC Service Controls.
- β’ Cap what any project may do with Organization Policy constraints, rolled out in dry-run first.
- β’ Control encryption with CMEK, and treat the key as a live dependency rather than a checkbox.
- β’ Serve the planet from one anycast IP, and diagnose a 502 from the health-check firewall rule.
- β’ Combine Private Google Access, Cloud NAT, PSC and VPC Service Controls without confusing them.
- β’ Make Cloud Run cheap by setting concurrency, and roll back a revision in seconds.
- β’ Separate Cloud SQL HA, read replicas and PITR β three features solving three problems.
- β’ Manage GKE upgrades with channels and windows, and unblock a drain stalled by a PDB.
- β’ Classify every resource as zonal, regional or multi-regional and name what dies with it.
- β’ Tell an IAM denial from an org policy, a VPC-SC perimeter and a disabled service.
- β’ Ship a landing zone in Terraform and prove it with seven drills, not a screenshot.
- β’ Version, rotate and audit secrets in Secret Manager, and stop the rotation that breaks callers.
- β’ Model DNS with public, private and peered zones, and read a resolution failure end to end.
- β’ Reserve and attach static IPs correctly, and expose services with the Gateway API.
- β’ Build reproducible VM images and treat instances as disposable rather than as pets.
- β’ Judge honestly whether AlloyDB earns its price over Cloud SQL, using your own measurements.
- β’ Turn bucket changes into events with Pub/Sub notifications or Eventarc, without a self-triggering loop.
- β’ Place logs in the right bucket with the right retention, and query them with SQL in Log Analytics.
- β’ Route logs to buckets, BigQuery, GCS and Pub/Sub, and fix the writer identity that fails silently.
- β’ Set ack deadlines, ordering keys and dead-letter topics from measured consumer behaviour.
- β’ Replicate a database with Datastream CDC, and alert on the replication slot before the disk fills.
π‘οΈ Best Practices in Production
The short version of this path. Every lesson also ends with the specific mistake it exists to prevent.
- β Design the resource hierarchy first β org, folders, projects β because IAM inherits downward.
- β Use the project as the blast radius and billing boundary, one per workload and environment.
- β Prefer Workload Identity Federation and service-account impersonation over downloaded key files.
- β Grant predefined roles at the narrowest scope that works, and review with the Policy Analyzer.
- β Use Shared VPC with host and service projects rather than peering everything together.
- β Size GKE subnets with secondary ranges for pods and services before creating the cluster.
- β Route logs with sinks to a dedicated project, and set retention deliberately.
- β Apply labels consistently β they are how cost attribution and inventory work later.
- β Downloading service-account JSON keys. They do not expire and they leak.
- β Granting
roles/editorbecause it is quick β it is close to project admin. - β Creating a GKE cluster with default node pool settings, then discovering the IP range is too small.
- β Firewall rules with
0.0.0.0/0on the default network, which many projects still carry. - β Assuming Cloud Build's default pool can reach private resources β it cannot without a private pool.
- β Leaving budget alerts unset on a project, so the first signal is the invoice.
πΌ Interview Readiness
Once you reach the end of this path, test your engineering knowledge against real questions asked by top technical teams:
Design a systematic inventory-building process for a GCP account with no existing documentation, using the gcloud CLI. What categories would you capture, and why does the organizational structure matter?
Cloud Cost OptimizationWhy is GCP's Active Assist described as a meaningful differentiator versus AWS Trusted Advisor?
Cloud Cost Optimization