πŸš€
Active Module Path

GCP Operations

Master GCP operations: Resource Hierarchy & Labels, IAM & Service Account delegation, Shared VPCs & Firewall Tags, GCE MIGs & SUD/CUD cost optimization, GCS WORM Lifecycles, GKE Pod IP & Workload Identity, Cloud Logging Sinks, and Pub/Sub & BigQuery analytics.

28 lessons 649 min syllabus

Stage 1 β€” Hierarchy & Security

5 lessons
Lesson 1 β€’ ⏱️ 20m
GCP Resource Hierarchy, Projects & Labels vs Network Tags

How GCP structures resources through Organizations, Folders, and Projects, and why understanding the difference between Labels and Network Tags prevents severe security and billing mistakes.

βœ“
Lesson 2 β€’ ⏱️ 22m
Cloud IAM, Service Accounts & Credential Delegation

Master GCP Cloud IAM bindings, Primitive vs Predefined roles, Service Account security, and eliminating exported JSON keys with Service Account Impersonation.

βœ“
Lesson 3 β€’ ⏱️ 22m
Organization Policy & Guardrails

Constraints that cap what any project may do, why an org policy grants nothing, and how to roll one out in dry-run before it blocks a production deploy.

βœ“
Lesson 4 β€’ ⏱️ 22m
Cloud KMS, CMEK and Secret Manager

The three levels of key control, what CMEK actually changes about a managed service, and why the key becomes a dependency you have to operate.

βœ“
Lesson 5 β€’ ⏱️ 20m
Secret Manager in Depth

Versions as immutable payloads, why disabling comes before destroying, replication and residency, rotation notifications, and consuming a secret with no key file anywhere.

βœ“

Stage 2 β€” Networking & Shared Topologies

5 lessons
Lesson 6 β€’ ⏱️ 25m
Global VPC, Subnets, Firewall Rules & Shared VPC Architecture

Understand GCP's unique global VPC network design, regional subnets, firewall rule targeting via Network Tags, and enterprise Shared VPC topologies.

βœ“
Lesson 7 β€’ ⏱️ 24m
Cloud Load Balancing, CDN and Cloud Armor

One anycast IP for the planet, the five objects between it and your backends, and the health-check firewall rule that causes most 502s.

βœ“
Lesson 8 β€’ ⏱️ 24m
Private Connectivity and Service Perimeters

Private Google Access, Cloud NAT, Private Service Connect and VPC Service Controls β€” four mechanisms that solve four different problems and are constantly confused.

βœ“
Lesson 9 β€’ ⏱️ 20m
Cloud DNS: Zones, Records and Routing

Public, private, forwarding and peering zones, the record types you will actually write, why you cannot CNAME an apex, and TTL as your rollback speed.

βœ“
Lesson 10 β€’ ⏱️ 20m
IP Addressing and the Gateway API

Ephemeral versus reserved, regional versus global, why an apex record forces you to reserve, and giving a GKE Gateway a stable address.

βœ“

Stage 3 β€” Compute & Storage Operations

7 lessons
Lesson 11 β€’ ⏱️ 25m
Compute Engine (GCE), MIGs, Auto-healing & Cost Optimization (SUD/CUD)

Master GCE machine families, Custom machine types, Managed Instance Groups (MIGs), auto-healing, Spot VMs, and automated vs committed cost optimization.

βœ“
Lesson 12 β€’ ⏱️ 20m
Google Cloud Storage (GCS) & Lifecycle Policies

Master GCS object storage classes, automated lifecycle transitions, bucket retention locks for WORM compliance, and Uniform Bucket-Level Access.

βœ“
Lesson 13 β€’ ⏱️ 24m
Serverless: Cloud Run and Cloud Functions

Concurrency as the thing that makes Cloud Run cheap, cold starts and what actually fixes them, revisions and traffic splitting, and when a cluster is still the answer.

βœ“
Lesson 14 β€’ ⏱️ 24m
Cloud SQL and Managed Databases

HA versus read replicas versus backups β€” three features that solve three different problems β€” plus the Auth Proxy, maintenance windows, and choosing between Cloud SQL, AlloyDB, Spanner and Firestore.

βœ“
Lesson 15 β€’ ⏱️ 22m
Compute Engine VMs in Depth

Images, disks and snapshots, what metadata actually controls, OS Login and IAP instead of SSH keys, and what survives a stop, a delete and a live migration.

βœ“
Lesson 16 β€’ ⏱️ 22m
AlloyDB: PostgreSQL Beyond Cloud SQL

Compute separated from distributed storage, read pools that scale independently, the columnar engine, and an honest account of when Cloud SQL is still the right answer.

βœ“
Lesson 17 β€’ ⏱️ 20m
Bucket Notifications and Eventarc

Turning an object change into a message, why every consumer must be idempotent, the loop that costs money, and choosing between Pub/Sub notifications and Eventarc.

βœ“

Stage 4 β€” Kubernetes, Observability & Analytics

8 lessons
Lesson 18 β€’ ⏱️ 25m
GKE Operations: Control Plane, Node Pools & Pod IP Allocation

Master Google Kubernetes Engine (GKE) Autopilot vs Standard, VPC-Native Alias IP networking, Node Pool management, and pod identity delegation via Workload Identity.

βœ“
Lesson 19 β€’ ⏱️ 22m
Cloud Operations Suite: Logging, Sinks, Monitoring & Alerts

Master GCP Cloud Logging, Audit vs Data Access logs, building Log Router Sinks to BigQuery & GCS, log-based metrics, and Cloud Monitoring alert policies.

βœ“
Lesson 20 β€’ ⏱️ 25m
Managed Data Services, Pub/Sub, BigQuery & Cloud Security Controls

Master Cloud Pub/Sub vs Cloud Tasks messaging, BigQuery partitioning/clustering optimization, managed databases, VPC Service Controls, and zero-trust IAP access.

βœ“
Lesson 21 β€’ ⏱️ 24m
GKE Day-2: Upgrades, Channels and Cost

Upgrades happen whether you plan them or not β€” release channels, maintenance windows, surge behaviour, and the Autopilot-versus-Standard cost question answered with numbers.

βœ“
Lesson 22 β€’ ⏱️ 22m
Cloud Logging In Depth: Buckets, Views and Queries

Where a log entry physically lives, log buckets and the _Default/_Required split, analytics views, the query language worth learning properly, and how retention becomes a bill.

βœ“
Lesson 23 β€’ ⏱️ 20m
The Log Router: Sinks, Filters and Exports

How routing actually evaluates, the four destinations and what each is for, aggregated sinks at the organisation level, and the writer identity that makes exports silently fail.

βœ“
Lesson 24 β€’ ⏱️ 22m
Pub/Sub In Depth: Delivery, Ordering and Backlogs

Push versus pull, ack deadlines and the redelivery loop, ordering keys and what they cost, dead-letter topics, exactly-once, and the one metric that tells you a consumer is broken.

βœ“
Lesson 25 β€’ ⏱️ 22m
Datastream and Change Data Capture

Reading the database's replication log instead of querying it, the backfill-then-stream model, connectivity choices, schema drift, and when CDC is the wrong answer.

βœ“

Stage 5 β€” Reliability & Capstone

3 lessons
Lesson 26 β€’ ⏱️ 24m
Resilience and Disaster Recovery on GCP

Zonal, regional and multi-regional resource scope, turning availability into an RTO and an RPO, and why a global load balancer makes regional failover a health-check outcome.

βœ“
Lesson 27 β€’ ⏱️ 22m
The GCP Failure Playbook

A six-step sweep for any GCP incident, telling an IAM denial from an org policy denial, and the quota and propagation failures that present as random.

βœ“
Lesson 28 β€’ ⏱️ 45m
Project: A Production Landing Zone on GCP

Build the whole module as one environment β€” org policies, Shared VPC, private workloads, CMEK, a global load balancer β€” then break it seven ways and prove it recovers.

βœ“

πŸ—ΊοΈ Beginner β†’ Expert Roadmap

5 stages with prerequisites and a concrete mastery check at each.

→

🎯 What You'll Learn

  • β€’ Organize GCP infrastructure with Organizations, Folders, Projects, and differentiate Labels from Network Tags.
  • β€’ Implement Cloud IAM least privilege, predefined vs primitive roles, and Service Account Impersonation without long-lived keys.
  • β€’ Design Global VPC networks, Regional Subnets, Firewall rules with Network Tags, and Shared VPC topologies.
  • β€’ Optimize GCE machine families, Custom machine types, Spot VMs, Sustained Use Discounts (SUD), and Committed Use Discounts (CUD).
  • β€’ Configure GCS storage classes (Standard/Nearline/Coldline/Archive), Bucket Lifecycle policies, and immutable Bucket Locks.
  • β€’ Deploy GKE clusters with VPC-Native Alias IP networking, custom Node Pools, and pod-level IAM via Workload Identity.
  • β€’ Build Cloud Logging Sinks to GCS, BigQuery, and Pub/Sub, configure log-based metrics, and set up Cloud Monitoring alerts.
  • β€’ Architect event-driven pipelines with Pub/Sub vs Cloud Tasks, BigQuery partitioning/clustering, and zero-trust IAP/VPC Service Controls.
  • β€’ Cap what any project may do with Organization Policy constraints, rolled out in dry-run first.
  • β€’ Control encryption with CMEK, and treat the key as a live dependency rather than a checkbox.
  • β€’ Serve the planet from one anycast IP, and diagnose a 502 from the health-check firewall rule.
  • β€’ Combine Private Google Access, Cloud NAT, PSC and VPC Service Controls without confusing them.
  • β€’ Make Cloud Run cheap by setting concurrency, and roll back a revision in seconds.
  • β€’ Separate Cloud SQL HA, read replicas and PITR β€” three features solving three problems.
  • β€’ Manage GKE upgrades with channels and windows, and unblock a drain stalled by a PDB.
  • β€’ Classify every resource as zonal, regional or multi-regional and name what dies with it.
  • β€’ Tell an IAM denial from an org policy, a VPC-SC perimeter and a disabled service.
  • β€’ Ship a landing zone in Terraform and prove it with seven drills, not a screenshot.
  • β€’ Version, rotate and audit secrets in Secret Manager, and stop the rotation that breaks callers.
  • β€’ Model DNS with public, private and peered zones, and read a resolution failure end to end.
  • β€’ Reserve and attach static IPs correctly, and expose services with the Gateway API.
  • β€’ Build reproducible VM images and treat instances as disposable rather than as pets.
  • β€’ Judge honestly whether AlloyDB earns its price over Cloud SQL, using your own measurements.
  • β€’ Turn bucket changes into events with Pub/Sub notifications or Eventarc, without a self-triggering loop.
  • β€’ Place logs in the right bucket with the right retention, and query them with SQL in Log Analytics.
  • β€’ Route logs to buckets, BigQuery, GCS and Pub/Sub, and fix the writer identity that fails silently.
  • β€’ Set ack deadlines, ordering keys and dead-letter topics from measured consumer behaviour.
  • β€’ Replicate a database with Datastream CDC, and alert on the replication slot before the disk fills.

πŸ›‘οΈ Best Practices in Production

The short version of this path. Every lesson also ends with the specific mistake it exists to prevent.

Do this
  • βœ“ Design the resource hierarchy first β€” org, folders, projects β€” because IAM inherits downward.
  • βœ“ Use the project as the blast radius and billing boundary, one per workload and environment.
  • βœ“ Prefer Workload Identity Federation and service-account impersonation over downloaded key files.
  • βœ“ Grant predefined roles at the narrowest scope that works, and review with the Policy Analyzer.
  • βœ“ Use Shared VPC with host and service projects rather than peering everything together.
  • βœ“ Size GKE subnets with secondary ranges for pods and services before creating the cluster.
  • βœ“ Route logs with sinks to a dedicated project, and set retention deliberately.
  • βœ“ Apply labels consistently β€” they are how cost attribution and inventory work later.
Avoid this
  • βœ— Downloading service-account JSON keys. They do not expire and they leak.
  • βœ— Granting roles/editor because it is quick β€” it is close to project admin.
  • βœ— Creating a GKE cluster with default node pool settings, then discovering the IP range is too small.
  • βœ— Firewall rules with 0.0.0.0/0 on the default network, which many projects still carry.
  • βœ— Assuming Cloud Build's default pool can reach private resources β€” it cannot without a private pool.
  • βœ— Leaving budget alerts unset on a project, so the first signal is the invoice.