Google Kubernetes Engine (GKE) is Google’s managed Kubernetes service. Because Kubernetes originated inside Google (from Borg), GKE features deep, native integration into GCP’s networking, IAM, and observability stacks.
Topic 1: GKE Standard vs. Autopilot
GCP offers two operational modes for GKE:
| Feature | GKE Standard | GKE Autopilot |
|---|---|---|
| Node Management | You configure and manage Node Pools (machine types, disk sizes, OS images) | Fully managed by Google — no node management or SSH |
| Billing Model | Pay per GCE VM node provisioned (regardless of pod utilization) | Pay strictly per requested Pod CPU, Memory, and Storage |
| Customization | Full control over daemonsets, kernel parameters, custom node pools | Enforces hardened security benchmarks (no privileged pods) |
| Autoscaling | Cluster Autoscaler + HPA / VPA | Fully automated pod-driven scaling |
Topic 2: VPC-Native Clusters & Alias IP Networking
In legacy Kubernetes setups (Routes-based), pod traffic requires overlay encapsulation (Flannel/VXLAN) or complex host routing.
GKE Production Standard: VPC-Native Clusters: VPC-Native clusters assign real, routable internal IP addresses to every Pod directly from the VPC subnet’s Secondary IPv4 CIDR range (Alias IPs):
- Primary Subnet Range (
10.100.0.0/20): Assigns internal IPs to GCE Node VMs. - Secondary Pod Range (
10.200.0.0/14): Assigns IP addresses directly to Pods. - Secondary Service Range (
10.204.0.0/20): Assigns ClusterIP addresses for K8s Services.
Benefits of VPC-Native Networking:
- Pods can communicate directly with Cloud SQL, Memorystore, and on-premises endpoints without NAT or gateway overlays.
- Better network throughput and lower latency.
- Network policies and VPC firewall rules can target Pod IP blocks directly.
Topic 3: Node Pool Architecture & Spot Node Pools
A Node Pool is a subset of worker nodes within a cluster that share identical machine configurations:
# Create a GKE VPC-Native Cluster with Workload Identity enabled
gcloud container clusters create prod-gke-cluster \
--region=us-central1 \
--enable-ip-alias \
--subnetwork=prod-subnet-us-central1 \
--cluster-secondary-range-name=gke-pods \
--services-secondary-range-name=gke-services \
--workload-pool=$(gcloud config get-value project).svc.id.goog
# Add a dedicated Spot VM Node Pool for fault-tolerant workloads
gcloud container node-pools create spot-pool \
--cluster=prod-gke-cluster \
--region=us-central1 \
--spot \
--machine-type=e2-standard-4 \
--enable-autoscaling \
--min-nodes=1 \
--max-nodes=10
Topic 4: Workload Identity (Pod-Level GCP IAM)
In standard Kubernetes, pods needing access to GCP APIs (GCS, BigQuery, Cloud SQL) often mounted exported JSON service account key files stored in Kubernetes Secrets. This is a severe security vulnerability.
GKE Workload Identity links a Kubernetes ServiceAccount (K8s SA) directly to a GCP IAM ServiceAccount (GCP SA):
- GKE runs a metadata server daemon on every node.
- When a container inside a Pod queries
http://169.254.169.254/computeMetadata/v1/instance/service-accounts/default/token, GKE’s metadata server intercepts the call. - GKE validates the Pod’s Kubernetes token, verifies the Workload Identity binding, and returns a short-lived GCP OAuth2 access token for the target GCP Service Account!
# Allow the K8s ServiceAccount 'sa-backend' in namespace 'production'
# to impersonate the GCP Service Account 'gcp-sa-backend'
gcloud iam service-accounts add-iam-policy-binding \
gcp-sa-backend@my-project-id.iam.gserviceaccount.com \
--role="roles/iam.workloadIdentityUser" \
--member="serviceAccount:my-project-id.svc.id.goog[production/sa-backend]"
Common mistake: Accepting the default secondary IP ranges when creating a VPC-native cluster. Pod and service ranges cannot be changed afterwards, and every pod consumes a real alias IP — so the cluster that sized comfortably for twenty nodes stops scheduling at sixty, and the fix is a new cluster rather than a setting.