Compute Engine (GCE), MIGs, Auto-healing & Cost Optimization (SUD/CUD)

Master GCE machine families, Custom machine types, Managed Instance Groups (MIGs), auto-healing, Spot VMs, and automated vs committed cost optimization.

intermediate 25 min lesson hands-on task included

Compute Engine (GCE) provides virtual machines running on Google’s infrastructure. In production DevOps, standalone GCE VMs are rarely deployed manually — instead, workloads run inside Managed Instance Groups (MIGs) backed by Instance Templates and automatic cost optimization strategies.


Topic 1: Machine Families & Custom Machine Types

COMPUTE ENGINE (GCE) COST OPTIMIZATION MATRIX 1. MACHINE FAMILY SELECTION & CUSTOM MACHINE TYPES E2 / N2 (General) Web apps & microservices Best price/perf ratio C2 / C3 (Compute) High CPU throughput Batch jobs & gaming M2 / M3 (Memory) Large DBs & SAP HANA Up to 12TB RAM Custom Types Specify exact vCPU/RAM Avoid overprovisioning 2. AUTOMATED VS COMMITTED VS INTERRUPTIBLE DISCOUNTS SUSTAINED USE (SUD) • Automatic up to 30% discount • Triggers when VM runs >25% of month • No upfront commitment required Applies to N1/N2/V100 VMs COMMITTED USE (CUD) • Up to 57% (Compute) / 70% (Memory) • 1-year or 3-year commitment • Resource-based or Spend-based For predictable baseline workloads SPOT / PREEMPTIBLE • 60% to 91% cost savings • 30-second shutdown notice • Ideal for fault-tolerant GKE/batch Use shutdown scripts for clean exit MANAGED INSTANCE GROUPS (MIG): Combine Instance Templates + Auto-healing health checks + Regional distribution for high availability.
GCE Cost Optimization Matrix: Machine Family selection, Custom Machine Types, Sustained Use Discounts (SUD), Committed Use Discounts (CUD), and Spot VMs.

GCP categorizes GCE VMs into workload-tailored machine families:

  • E2 / N2 / N2D (General Purpose): Balanced price-performance ratio for web applications, microservices, and small databases. E2 features dynamic cost optimization; N2D uses AMD EPYC processors (often 10–20% cheaper than Intel).
  • C2 / C3 (Compute Optimized): Ultra-high single-thread frequency for high-performance computing, intensive batch processing, and gaming servers.
  • M2 / M3 (Memory Optimized): Massive RAM footprints (up to 12 TB) designed for large in-memory databases like SAP HANA or giant Redis clusters.
  • Shared-Core (e2-micro, e2-small, e2-medium): Timeshares physical CPU cores; cost-effective for lightweight dev/staging services.

Custom Machine Types: Unlike AWS (where you must pick rigid instance sizes like c6i.xlarge), GCP allows you to specify exact vCPU and RAM numbers (e.g., a custom VM with 6 vCPUs and 22 GB RAM):

# Create a Custom Machine Type with 6 vCPUs and 22 GB memory
gcloud compute instances create custom-app-vm \
  --zone=us-central1-a \
  --custom-cpu=6 \
  --custom-memory=22GiB \
  --network-interface=subnet=prod-subnet-us-central1

Topic 2: Managed Instance Groups (MIGs) & Auto-Healing

A Managed Instance Group (MIG) operates identical VMs instantiated from an Instance Template:

  1. Regional MIGs (High Availability): Distributes VM instances across 3 Availability Zones in a region. If an entire zone fails, the MIG automatically recreates missing capacity in the remaining zones.
  2. Immutable Instance Templates: Instance templates cannot be modified after creation. To roll out an update, create template-v2 and initiate a MIG Rolling Update (canary or blue-green replacement).
  3. Auto-Healing: You attach an HTTP/TCP health check to the MIG. If a VM fails the health check criteria (e.g., 3 consecutive HTTP 500s), the MIG automatically deletes the degraded VM and provisions a fresh replacement VM.
# 1. Create an immutable Instance Template
gcloud compute instance-templates create app-template-v1 \
  --machine-type=e2-standard-2 \
  --subnet=prod-subnet-us-central1 \
  --tags=target-web-server \
  --metadata=startup-script="#!/bin/bash
apt-get update && apt-get install -y nginx"

# 2. Create an Auto-Healing Health Check
gcloud compute health-checks create http app-health-check \
  --port=80 \
  --request-path=/healthz

# 3. Create a Regional MIG with Auto-Healing
gcloud compute instance-groups managed create app-mig-prod \
  --template=app-template-v1 \
  --size=3 \
  --region=us-central1 \
  --health-check=app-health-check \
  --initial-delay=120

Topic 3: GCP Cost Optimization: SUD vs. CUD vs. Spot VMs

GCP offers three major discount mechanisms that drastically reduce cloud spend:

1. Sustained Use Discounts (SUD) — Automated

  • No contract or configuration required. Google automatically applies discounts of up to 30% on GCE instance families (N1, N2) when a VM runs for more than 25% of a billing month.

2. Committed Use Discounts (CUD) — Contractual

  • Commit to a baseline level of compute (vCPUs/RAM or Spend) for 1 year (up to 37% off) or 3 years (up to 57–70% off).
  • Resource-Based CUD: Bound to specific instance families in a specific region.
  • Flexible / Spend-Based CUD: Applies across multiple instance families and regions based on dollars per hour committed.

3. Spot VMs (formerly Preemptible VMs) — Interruptible

  • Excess GCP compute capacity sold at 60% to 91% discount.
  • GCP can reclaim (terminate) a Spot VM at any time with a 30-second shutdown notification.
  • Perfect for fault-tolerant GKE node pools, batch data processing, and CI/CD runners. Use graceful shutdown scripts to handle the 30-second signal (SIGTERM).

Common mistake: Buying committed use discounts before the workload has stopped moving. A CUD is a one to three year commitment to a machine family in a region — right-size and migrate architecture first, then commit to the floor that remains. Sustained use discounts apply automatically and need no decision at all.