๐Ÿš€
OPERATIONAL ROADMAP ยท 4 STAGES ยท 8 LESSONS

GCP Operations Mastery Path

From organizational resource hierarchy down to GKE Workload Identity, Shared VPC topologies, Log Router Sinks, and BigQuery analytics โ€” master Google Cloud Platform the way it operates under production load.

8 Production Lessons
~10.8h Total Duration
8 Hands-on Tasks
STAGE 1

Establish organizational boundaries, IAM least privilege & Service Account delegation

5 lessons ยท ~2h
PREREQUISITE

Basic understanding of Linux administration and cloud access control concepts.

OPERATIONAL FOCUS

The GCP Resource Hierarchy (Org -> Folders -> Projects) dictates permission inheritance across all services.

STAGE CAPABILITIES YOU WILL MASTER

  • โœ“ Structure GCP environments using Organizations, Folders, and Projects
  • โœ“ Distinguish resource Labels (billing/cost attribution) from Network Tags (VPC routing & firewalls)
  • โœ“ Apply Cloud IAM least-privilege principles using Predefined and Custom roles
  • โœ“ Eliminate long-lived JSON service account keys with Service Account Impersonation
  • โœ“ Cap what any project may do with Organization Policy constraints โ€” rolled out in dry-run first
  • โœ“ Control encryption with Cloud KMS and CMEK, and keep credentials in Secret Manager
  • โœ“ Version, rotate and audit secrets in Secret Manager without breaking the callers that read them
STAGE 2

Master Global VPCs, Regional Subnets, Firewall Network Tags & Shared VPCs

5 lessons ยท ~2h
PREREQUISITE

Completion of Stage 1 and familiarity with CIDR notation and VPC routing.

OPERATIONAL FOCUS

GCP VPCs are global resources containing regional subnets that span all availability zones in that region.

STAGE CAPABILITIES YOU WILL MASTER

  • โœ“ Design Global VPCs with Custom Subnets and alias IP allocations
  • โœ“ Apply stateful VPC firewall rules using Network Tags and Service Accounts
  • โœ“ Architect Shared VPCs separating Host Projects (network) from Service Projects (workloads)
  • โœ“ Evaluate trade-offs between Shared VPC, VPC Network Peering, and Cloud Interconnect/VPN
  • โœ“ Serve the planet from one anycast IP with the global load balancer, Cloud CDN and Cloud Armor
  • โœ“ Combine Private Google Access, Cloud NAT, Private Service Connect and VPC Service Controls
  • โœ“ Model DNS with public, private and peered zones, and trace a resolution failure end to end
  • โœ“ Reserve and attach static IP addresses correctly, and expose services with the Gateway API
STAGE 3

Operate GCE MIGs, auto-healing, GCS WORM Lifecycles & CUD/SUD cost optimization

7 lessons ยท ~2.5h
PREREQUISITE

Completion of Stage 2 and understanding of virtual machine life-cycles and object storage.

OPERATIONAL FOCUS

Cost optimization on GCP relies on automated Sustained Use Discounts (SUD) and targeted Committed Use Discounts (CUD).

4

Compute Engine (GCE), MIGs, Auto-healing & Cost Optimization (SUD/CUD)

Master GCE machine families, Custom machine types, Managed Instance Groups (MIGs), auto-healing, Spot VMs, and automated vs committed cost optimization.

25m Read โ†’
5

Google Cloud Storage (GCS) & Lifecycle Policies

Master GCS object storage classes, automated lifecycle transitions, bucket retention locks for WORM compliance, and Uniform Bucket-Level Access.

20m Read โ†’
13

Serverless: Cloud Run and Cloud Functions

Concurrency as the thing that makes Cloud Run cheap, cold starts and what actually fixes them, revisions and traffic splitting, and when a cluster is still the answer.

24m Read โ†’
14

Cloud SQL and Managed Databases

HA versus read replicas versus backups โ€” three features that solve three different problems โ€” plus the Auth Proxy, maintenance windows, and choosing between Cloud SQL, AlloyDB, Spanner and Firestore.

24m Read โ†’
22

Compute Engine VMs in Depth

Images, disks and snapshots, what metadata actually controls, OS Login and IAP instead of SSH keys, and what survives a stop, a delete and a live migration.

22m Read โ†’
23

AlloyDB: PostgreSQL Beyond Cloud SQL

Compute separated from distributed storage, read pools that scale independently, the columnar engine, and an honest account of when Cloud SQL is still the right answer.

22m Read โ†’
24

Bucket Notifications and Eventarc

Turning an object change into a message, why every consumer must be idempotent, the loop that costs money, and choosing between Pub/Sub notifications and Eventarc.

20m Read โ†’

STAGE CAPABILITIES YOU WILL MASTER

  • โœ“ Select optimal GCE machine families (N2, C2, E2) and create Custom Machine Types
  • โœ“ Configure Managed Instance Groups (MIGs) with immutable templates, auto-healing & canary updates
  • โœ“ Leverage Spot/Preemptible VMs with 30-second shutdown notice handling
  • โœ“ Configure GCS storage classes, automated Bucket Lifecycle rules, and immutable Bucket Locks
  • โœ“ Run containers without a cluster on Cloud Run โ€” concurrency, cold starts and revision rollback
  • โœ“ Operate Cloud SQL knowing HA, read replicas and PITR solve three different problems
  • โœ“ Build reproducible VM images and treat instances as disposable rather than as pets
  • โœ“ Judge whether AlloyDB earns its price over Cloud SQL from your own measurements
  • โœ“ Turn bucket changes into events with Pub/Sub notifications or Eventarc, without a self-triggering loop
STAGE 4

Run GKE with Workload Identity, Cloud Logging Sinks & Zero-Trust IAP

8 lessons ยท ~3h
PREREQUISITE

Completion of Stage 3 and familiarity with Kubernetes and central log aggregation.

OPERATIONAL FOCUS

GKE VPC-Native clusters assign routable Pod IPs directly from subnet secondary CIDRs.

6

GKE Operations: Control Plane, Node Pools & Pod IP Allocation

Master Google Kubernetes Engine (GKE) Autopilot vs Standard, VPC-Native Alias IP networking, Node Pool management, and pod identity delegation via Workload Identity.

25m Read โ†’
7

Cloud Operations Suite: Logging, Sinks, Monitoring & Alerts

Master GCP Cloud Logging, Audit vs Data Access logs, building Log Router Sinks to BigQuery & GCS, log-based metrics, and Cloud Monitoring alert policies.

22m Read โ†’
8

Managed Data Services, Pub/Sub, BigQuery & Cloud Security Controls

Master Cloud Pub/Sub vs Cloud Tasks messaging, BigQuery partitioning/clustering optimization, managed databases, VPC Service Controls, and zero-trust IAP access.

25m Read โ†’
15

GKE Day-2: Upgrades, Channels and Cost

Upgrades happen whether you plan them or not โ€” release channels, maintenance windows, surge behaviour, and the Autopilot-versus-Standard cost question answered with numbers.

24m Read โ†’
25

Cloud Logging In Depth: Buckets, Views and Queries

Where a log entry physically lives, log buckets and the _Default/_Required split, analytics views, the query language worth learning properly, and how retention becomes a bill.

22m Read โ†’
26

The Log Router: Sinks, Filters and Exports

How routing actually evaluates, the four destinations and what each is for, aggregated sinks at the organisation level, and the writer identity that makes exports silently fail.

20m Read โ†’
27

Pub/Sub In Depth: Delivery, Ordering and Backlogs

Push versus pull, ack deadlines and the redelivery loop, ordering keys and what they cost, dead-letter topics, exactly-once, and the one metric that tells you a consumer is broken.

22m Read โ†’
28

Datastream and Change Data Capture

Reading the database's replication log instead of querying it, the backfill-then-stream model, connectivity choices, schema drift, and when CDC is the wrong answer.

22m Read โ†’

STAGE CAPABILITIES YOU WILL MASTER

  • โœ“ Operate GKE Autopilot & Standard clusters with Alias IP networking and custom Node Pools
  • โœ“ Secure pod-to-GCP communications using Workload Identity (K8s SA to GCP SA mapping)
  • โœ“ Configure Cloud Logging Router Sinks exporting logs to BigQuery, GCS, and Pub/Sub
  • โœ“ Implement zero-trust security perimeters using VPC Service Controls, Cloud Armor, and IAP
  • โœ“ Manage GKE upgrades with release channels, maintenance windows and surge settings
  • โœ“ Decide Autopilot versus Standard from measured cost rather than from preference
  • โœ“ Place logs in the right bucket with the right retention, and query them with SQL in Log Analytics
  • โœ“ Route logs to buckets, BigQuery, GCS and Pub/Sub, and fix the writer identity that fails silently
  • โœ“ Set Pub/Sub ack deadlines, ordering keys and dead-letter topics from measured consumer behaviour
  • โœ“ Replicate a database with Datastream CDC and alert on the replication slot before the disk fills
STAGE 5

Survive a zone, diagnose a denial, and prove a landing zone with drills

3 lessons ยท ~1.5h
PREREQUISITE

Stages 1-4. The capstone project assumes every one of them.

OPERATIONAL FOCUS

Everything here is about failure: what breaks with a zone, what a denial actually came from, and what a drill proves that a diagram cannot.

STAGE CAPABILITIES YOU WILL MASTER

  • โœ“ Classify every resource as zonal, regional or multi-regional and name what dies with it
  • โœ“ Turn an availability requirement into a written RTO and RPO, then pick the cheapest strategy that meets both
  • โœ“ Run a fixed six-step triage sweep instead of debugging by intuition
  • โœ“ Tell an IAM denial from an org policy, a VPC-SC perimeter and a disabled service
  • โœ“ Find the quota that would break the workload first, before traffic does
  • โœ“ Ship a landing zone in Terraform and prove it with seven drills