Compute Engine VMs in Depth

Images, disks and snapshots, what metadata actually controls, OS Login and IAP instead of SSH keys, and what survives a stop, a delete and a live migration.

intermediate 22 min lesson hands-on task included

A VM is a machine type, a disk created from an image, and a bag of metadata. Almost every Compute Engine surprise comes from one of those three.


Topic 1: The Anatomy

A VM IS A MACHINE TYPE, A DISK FROM AN IMAGE, AND SOME METADATA image public, custom, or a machine image boot disk pd-balanced by default; zonal the VM machine type + metadata + tags snapshot incremental, and REGIONAL METADATA IS THE CONFIGURATION SURFACE startup-script runs as root on every boot shutdown-script best effort, bounded enable-oslogin IAM-managed SSH, auditable 169.254.169.254 — the metadata server, same as every cloud WHAT SURVIVES WHAT stop/start → local SSD gone, ephemeral IP changes delete → boot disk gone unless --no-boot-disk-auto-delete live migration → maintenance without a reboot, by default Zonal disk = zonal VM. Regional PD replicates across two zones. OS LOGIN + IAP REPLACES SSH KEYS AND BASTIONS ENTIRELY gcloud compute ssh vm-1 --tunnel-through-iap — no external IP, no port 22 open, every session in the audit log Bake images with a pipeline (Packer or Cloud Build), reference them by family, and let a MIG replace instances rather than patching them in place.
Image to disk to VM, with a snapshot as the regional escape hatch. The right-hand panel is the list of what survives which operation — the source of most Compute Engine surprises.
gcloud compute instances create app-1 \
  --zone=europe-west1-b \
  --machine-type=e2-standard-4 \
  --image-family=debian-12 --image-project=debian-cloud \
  --boot-disk-type=pd-balanced --boot-disk-size=50GB \
  --no-address \
  --metadata=enable-oslogin=TRUE \
  --metadata-from-file=startup-script=./startup.sh \
  --service-account=app@acme.iam.gserviceaccount.com \
  --scopes=cloud-platform \
  --tags=http-backend \
  --shielded-secure-boot --shielded-vtpm --shielded-integrity-monitoring

Every flag there is a decision worth understanding:

--image-family tracks the latest image in a family rather than pinning a dated name, so a rebuild picks up patches. For reproducibility, pin the exact image and bump it deliberately — the same argument as container tags.

--no-address means no external IP. With OS Login and IAP this is entirely workable, and it should be the default.

--scopes=cloud-platform with a narrowly-scoped service account is the modern pattern. Legacy per-API scopes are a second, confusing authorisation layer on top of IAM; give the VM one service account with only the roles it needs and let IAM decide.

Shielded VM flags cost nothing and give you secure boot, a virtual TPM and integrity monitoring. There is an org policy to require them.


Topic 2: Disks and What Survives What

Persistent DiskLocal SSD
Survives stop/startYesNo — data is gone
Survives deleteOnly if auto-delete is offNo
Attachable elsewhereYesNever
Snapshot-ableYesNo
PerformanceScales with size and typeHighest available

Disk types, in the order you will consider them:

pd-balanced    the sensible default — SSD-backed, cheaper than pd-ssd
pd-ssd         when you need the IOPS and have measured it
pd-standard    HDD, for throughput-oriented bulk data
hyperdisk      newest, IOPS and throughput provisioned independently of size

Persistent disk performance scales with provisioned size — a 20 GB pd-balanced is slow because it is small, not because the type is wrong. That is the most common “the disk is slow” cause, and the fix is a bigger disk or Hyperdisk with explicit IOPS.

Zonal versus regional: a regional persistent disk replicates synchronously across two zones in the region and can be force-attached in the surviving zone. It costs roughly double and is the difference between a zonal VM and one that can be recovered elsewhere — the resilience lesson’s point, made concrete.

gcloud compute disks create data-1 --region=europe-west1 \
  --replica-zones=europe-west1-b,europe-west1-c --size=200GB --type=pd-balanced

Resize is online and one-way: you can grow a disk (and then the filesystem) without downtime; you cannot shrink it.


Topic 3: Snapshots, Images and Machine Images

Three different artifacts that people conflate:

Snapshot — an incremental, regional backup of a disk. The first is full, later ones store only changed blocks, and each is independently restorable. This is your cross-zone escape hatch: a zonal disk plus a snapshot equals a disk in any zone.

gcloud compute resource-policies create snapshot-schedule daily-snap \
  --region=europe-west1 --max-retention-days=14 \
  --daily-schedule --start-time=02:00 --storage-location=eu

gcloud compute disks add-resource-policies data-1 --zone=europe-west1-b \
  --resource-policies=daily-snap

Custom image — a bootable template for new VMs. This is what an image pipeline produces, and what a MIG’s instance template references.

Machine image — the whole VM: all disks, metadata, and configuration. Convenient for lift-and-shift and for cloning a snowflake you have not finished taming.

Snapshots are crash-consistent, not application-consistent. A snapshot of a running database is equivalent to pulling the power cord — usually recoverable, and not a backup strategy. Use the database’s own backup mechanism, or quiesce the filesystem for the instant of the snapshot.


Topic 4: Metadata, Startup Scripts and OS Login

The metadata server at 169.254.169.254 is how a VM learns about itself and gets credentials:

curl -H "Metadata-Flavor: Google" \
  http://169.254.169.254/computeMetadata/v1/instance/service-accounts/default/token

Metadata-Flavor: Google is required, which is GCP’s equivalent of the IMDSv2 protection from the AWS module — a plain SSRF that only issues a bare GET cannot read it.

Startup scripts run as root on every boot, not just the first:

#!/bin/bash
set -euo pipefail
exec > >(tee /var/log/startup.log|logger -t startup -s 2>/dev/console) 2>&1
apt-get update && apt-get install -y nginx
gcloud secrets versions access latest --secret=app-config > /etc/app/config.json
systemctl enable --now nginx

Two things to know: output goes to the serial console and the log file you redirect it to, and a failing startup script does not fail the instance — the VM reaches RUNNING, joins the load balancer and serves errors. Anything essential must be reflected in a health check.

OS Login replaces SSH key metadata with IAM:

gcloud compute instances add-metadata app-1 --metadata enable-oslogin=TRUE
gcloud compute instances add-iam-policy-binding app-1 --zone=europe-west1-b \
  --member=user:alice@acme.example --role=roles/compute.osLogin

Access is granted and revoked in IAM, sessions are audited, and there are no keys spread through instance metadata. Combined with IAP TCP forwarding, a VM needs no external IP and no open port 22:

gcloud compute ssh app-1 --zone=europe-west1-b --tunnel-through-iap

That pair is the GCP equivalent of AWS Session Manager, and it should be the default access path.


Topic 5: Availability Behaviour

Live migration moves a running VM to another host for maintenance with no reboot — the default, and one of the genuinely nice things about Compute Engine. Set --maintenance-policy=TERMINATE only when you must (GPUs, sole-tenant nodes with licensing constraints).

Automatic restart brings a VM back after a host failure, on by default.

Host maintenance events are visible in metadata, so a workload that needs to drain can react:

curl -H "Metadata-Flavor: Google" \
  http://169.254.169.254/computeMetadata/v1/instance/maintenance-event

Spot VMs are the same machine at 60–91% off, reclaimable at any time with a 30-second notice delivered as a shutdown script and a metadata change. They belong behind a MIG, for workloads that tolerate interruption — which is the cost lesson’s point applied here.

Sole-tenant nodes give you a physical host to yourself, for licensing or compliance. Expensive, and occasionally the only way to satisfy a rule.


Topic 6: Operating a Fleet Rather Than a Pet

The end state is that no VM is configured by hand:

□ an image pipeline (Packer or Cloud Build) producing versioned custom images
□ instance templates referencing an image, not a family, for production
□ MIGs with auto-healing health checks — the MIG lesson covers this
□ startup scripts limited to last-mile configuration
□ Ops Agent installed in the image for metrics and logs
□ OS Login + IAP; no SSH keys in metadata, no external IPs
□ snapshot schedules on every disk holding state
□ labels on every instance for cost attribution

The Ops Agent is worth calling out: without it you get no memory, disk or process metrics, because the hypervisor cannot see inside the guest. “Memory usage is missing from Cloud Monitoring” is always this.

curl -sSO https://dl.google.com/cloudagents/add-google-cloud-ops-agent-repo.sh
sudo bash add-google-cloud-ops-agent-repo.sh --also-install

Try it yourself: create a VM with a local SSD, write a file to it, stop and start the instance, and look for the file. It is gone, with no warning at any point in the process — which is the single most surprising Compute Engine behaviour and the reason local SSD belongs to caches and scratch space only.

Common mistake: treating a VM you configured by hand as the source of truth, and taking a machine image of it as the “backup”. The image captures the state including the drift, nobody can reproduce how it got there, and the next person to need a second instance clones a snowflake. Build the image from a pipeline, keep the inputs in Git, and let the VM be disposable.