GCP Operations Mastery Path
From organizational resource hierarchy down to GKE Workload Identity, Shared VPC topologies, Log Router Sinks, and BigQuery analytics โ master Google Cloud Platform the way it operates under production load.
Establish organizational boundaries, IAM least privilege & Service Account delegation
Basic understanding of Linux administration and cloud access control concepts.
The GCP Resource Hierarchy (Org -> Folders -> Projects) dictates permission inheritance across all services.
GCP Resource Hierarchy, Projects & Labels vs Network Tags
How GCP structures resources through Organizations, Folders, and Projects, and why understanding the difference between Labels and Network Tags prevents severe security and billing mistakes.
Cloud IAM, Service Accounts & Credential Delegation
Master GCP Cloud IAM bindings, Primitive vs Predefined roles, Service Account security, and eliminating exported JSON keys with Service Account Impersonation.
Organization Policy & Guardrails
Constraints that cap what any project may do, why an org policy grants nothing, and how to roll one out in dry-run before it blocks a production deploy.
Cloud KMS, CMEK and Secret Manager
The three levels of key control, what CMEK actually changes about a managed service, and why the key becomes a dependency you have to operate.
Secret Manager in Depth
Versions as immutable payloads, why disabling comes before destroying, replication and residency, rotation notifications, and consuming a secret with no key file anywhere.
STAGE CAPABILITIES YOU WILL MASTER
- โ Structure GCP environments using Organizations, Folders, and Projects
- โ Distinguish resource Labels (billing/cost attribution) from Network Tags (VPC routing & firewalls)
- โ Apply Cloud IAM least-privilege principles using Predefined and Custom roles
- โ Eliminate long-lived JSON service account keys with Service Account Impersonation
- โ Cap what any project may do with Organization Policy constraints โ rolled out in dry-run first
- โ Control encryption with Cloud KMS and CMEK, and keep credentials in Secret Manager
- โ Version, rotate and audit secrets in Secret Manager without breaking the callers that read them
Master Global VPCs, Regional Subnets, Firewall Network Tags & Shared VPCs
Completion of Stage 1 and familiarity with CIDR notation and VPC routing.
GCP VPCs are global resources containing regional subnets that span all availability zones in that region.
Global VPC, Subnets, Firewall Rules & Shared VPC Architecture
Understand GCP's unique global VPC network design, regional subnets, firewall rule targeting via Network Tags, and enterprise Shared VPC topologies.
Cloud Load Balancing, CDN and Cloud Armor
One anycast IP for the planet, the five objects between it and your backends, and the health-check firewall rule that causes most 502s.
Private Connectivity and Service Perimeters
Private Google Access, Cloud NAT, Private Service Connect and VPC Service Controls โ four mechanisms that solve four different problems and are constantly confused.
Cloud DNS: Zones, Records and Routing
Public, private, forwarding and peering zones, the record types you will actually write, why you cannot CNAME an apex, and TTL as your rollback speed.
IP Addressing and the Gateway API
Ephemeral versus reserved, regional versus global, why an apex record forces you to reserve, and giving a GKE Gateway a stable address.
STAGE CAPABILITIES YOU WILL MASTER
- โ Design Global VPCs with Custom Subnets and alias IP allocations
- โ Apply stateful VPC firewall rules using Network Tags and Service Accounts
- โ Architect Shared VPCs separating Host Projects (network) from Service Projects (workloads)
- โ Evaluate trade-offs between Shared VPC, VPC Network Peering, and Cloud Interconnect/VPN
- โ Serve the planet from one anycast IP with the global load balancer, Cloud CDN and Cloud Armor
- โ Combine Private Google Access, Cloud NAT, Private Service Connect and VPC Service Controls
- โ Model DNS with public, private and peered zones, and trace a resolution failure end to end
- โ Reserve and attach static IP addresses correctly, and expose services with the Gateway API
Operate GCE MIGs, auto-healing, GCS WORM Lifecycles & CUD/SUD cost optimization
Completion of Stage 2 and understanding of virtual machine life-cycles and object storage.
Cost optimization on GCP relies on automated Sustained Use Discounts (SUD) and targeted Committed Use Discounts (CUD).
Compute Engine (GCE), MIGs, Auto-healing & Cost Optimization (SUD/CUD)
Master GCE machine families, Custom machine types, Managed Instance Groups (MIGs), auto-healing, Spot VMs, and automated vs committed cost optimization.
Google Cloud Storage (GCS) & Lifecycle Policies
Master GCS object storage classes, automated lifecycle transitions, bucket retention locks for WORM compliance, and Uniform Bucket-Level Access.
Serverless: Cloud Run and Cloud Functions
Concurrency as the thing that makes Cloud Run cheap, cold starts and what actually fixes them, revisions and traffic splitting, and when a cluster is still the answer.
Cloud SQL and Managed Databases
HA versus read replicas versus backups โ three features that solve three different problems โ plus the Auth Proxy, maintenance windows, and choosing between Cloud SQL, AlloyDB, Spanner and Firestore.
Compute Engine VMs in Depth
Images, disks and snapshots, what metadata actually controls, OS Login and IAP instead of SSH keys, and what survives a stop, a delete and a live migration.
AlloyDB: PostgreSQL Beyond Cloud SQL
Compute separated from distributed storage, read pools that scale independently, the columnar engine, and an honest account of when Cloud SQL is still the right answer.
Bucket Notifications and Eventarc
Turning an object change into a message, why every consumer must be idempotent, the loop that costs money, and choosing between Pub/Sub notifications and Eventarc.
STAGE CAPABILITIES YOU WILL MASTER
- โ Select optimal GCE machine families (N2, C2, E2) and create Custom Machine Types
- โ Configure Managed Instance Groups (MIGs) with immutable templates, auto-healing & canary updates
- โ Leverage Spot/Preemptible VMs with 30-second shutdown notice handling
- โ Configure GCS storage classes, automated Bucket Lifecycle rules, and immutable Bucket Locks
- โ Run containers without a cluster on Cloud Run โ concurrency, cold starts and revision rollback
- โ Operate Cloud SQL knowing HA, read replicas and PITR solve three different problems
- โ Build reproducible VM images and treat instances as disposable rather than as pets
- โ Judge whether AlloyDB earns its price over Cloud SQL from your own measurements
- โ Turn bucket changes into events with Pub/Sub notifications or Eventarc, without a self-triggering loop
Run GKE with Workload Identity, Cloud Logging Sinks & Zero-Trust IAP
Completion of Stage 3 and familiarity with Kubernetes and central log aggregation.
GKE VPC-Native clusters assign routable Pod IPs directly from subnet secondary CIDRs.
GKE Operations: Control Plane, Node Pools & Pod IP Allocation
Master Google Kubernetes Engine (GKE) Autopilot vs Standard, VPC-Native Alias IP networking, Node Pool management, and pod identity delegation via Workload Identity.
Cloud Operations Suite: Logging, Sinks, Monitoring & Alerts
Master GCP Cloud Logging, Audit vs Data Access logs, building Log Router Sinks to BigQuery & GCS, log-based metrics, and Cloud Monitoring alert policies.
Managed Data Services, Pub/Sub, BigQuery & Cloud Security Controls
Master Cloud Pub/Sub vs Cloud Tasks messaging, BigQuery partitioning/clustering optimization, managed databases, VPC Service Controls, and zero-trust IAP access.
GKE Day-2: Upgrades, Channels and Cost
Upgrades happen whether you plan them or not โ release channels, maintenance windows, surge behaviour, and the Autopilot-versus-Standard cost question answered with numbers.
Cloud Logging In Depth: Buckets, Views and Queries
Where a log entry physically lives, log buckets and the _Default/_Required split, analytics views, the query language worth learning properly, and how retention becomes a bill.
The Log Router: Sinks, Filters and Exports
How routing actually evaluates, the four destinations and what each is for, aggregated sinks at the organisation level, and the writer identity that makes exports silently fail.
Pub/Sub In Depth: Delivery, Ordering and Backlogs
Push versus pull, ack deadlines and the redelivery loop, ordering keys and what they cost, dead-letter topics, exactly-once, and the one metric that tells you a consumer is broken.
Datastream and Change Data Capture
Reading the database's replication log instead of querying it, the backfill-then-stream model, connectivity choices, schema drift, and when CDC is the wrong answer.
STAGE CAPABILITIES YOU WILL MASTER
- โ Operate GKE Autopilot & Standard clusters with Alias IP networking and custom Node Pools
- โ Secure pod-to-GCP communications using Workload Identity (K8s SA to GCP SA mapping)
- โ Configure Cloud Logging Router Sinks exporting logs to BigQuery, GCS, and Pub/Sub
- โ Implement zero-trust security perimeters using VPC Service Controls, Cloud Armor, and IAP
- โ Manage GKE upgrades with release channels, maintenance windows and surge settings
- โ Decide Autopilot versus Standard from measured cost rather than from preference
- โ Place logs in the right bucket with the right retention, and query them with SQL in Log Analytics
- โ Route logs to buckets, BigQuery, GCS and Pub/Sub, and fix the writer identity that fails silently
- โ Set Pub/Sub ack deadlines, ordering keys and dead-letter topics from measured consumer behaviour
- โ Replicate a database with Datastream CDC and alert on the replication slot before the disk fills
Survive a zone, diagnose a denial, and prove a landing zone with drills
Stages 1-4. The capstone project assumes every one of them.
Everything here is about failure: what breaks with a zone, what a denial actually came from, and what a drill proves that a diagram cannot.
Resilience and Disaster Recovery on GCP
Zonal, regional and multi-regional resource scope, turning availability into an RTO and an RPO, and why a global load balancer makes regional failover a health-check outcome.
The GCP Failure Playbook
A six-step sweep for any GCP incident, telling an IAM denial from an org policy denial, and the quota and propagation failures that present as random.
Project: A Production Landing Zone on GCP
Build the whole module as one environment โ org policies, Shared VPC, private workloads, CMEK, a global load balancer โ then break it seven ways and prove it recovers.
STAGE CAPABILITIES YOU WILL MASTER
- โ Classify every resource as zonal, regional or multi-regional and name what dies with it
- โ Turn an availability requirement into a written RTO and RPO, then pick the cheapest strategy that meets both
- โ Run a fixed six-step triage sweep instead of debugging by intuition
- โ Tell an IAM denial from an org policy, a VPC-SC perimeter and a disabled service
- โ Find the quota that would break the workload first, before traffic does
- โ Ship a landing zone in Terraform and prove it with seven drills