A VM is a machine type, a disk created from an image, and a bag of metadata. Almost every Compute Engine surprise comes from one of those three.
Topic 1: The Anatomy
gcloud compute instances create app-1 \
--zone=europe-west1-b \
--machine-type=e2-standard-4 \
--image-family=debian-12 --image-project=debian-cloud \
--boot-disk-type=pd-balanced --boot-disk-size=50GB \
--no-address \
--metadata=enable-oslogin=TRUE \
--metadata-from-file=startup-script=./startup.sh \
--service-account=app@acme.iam.gserviceaccount.com \
--scopes=cloud-platform \
--tags=http-backend \
--shielded-secure-boot --shielded-vtpm --shielded-integrity-monitoring
Every flag there is a decision worth understanding:
--image-family tracks the latest image in a family rather than pinning a dated name, so a rebuild picks up patches. For reproducibility, pin the exact image and bump it deliberately — the same argument as container tags.
--no-address means no external IP. With OS Login and IAP this is entirely workable, and it should be the default.
--scopes=cloud-platform with a narrowly-scoped service account is the modern pattern. Legacy per-API scopes are a second, confusing authorisation layer on top of IAM; give the VM one service account with only the roles it needs and let IAM decide.
Shielded VM flags cost nothing and give you secure boot, a virtual TPM and integrity monitoring. There is an org policy to require them.
Topic 2: Disks and What Survives What
| Persistent Disk | Local SSD | |
|---|---|---|
| Survives stop/start | Yes | No — data is gone |
| Survives delete | Only if auto-delete is off | No |
| Attachable elsewhere | Yes | Never |
| Snapshot-able | Yes | No |
| Performance | Scales with size and type | Highest available |
Disk types, in the order you will consider them:
pd-balanced the sensible default — SSD-backed, cheaper than pd-ssd
pd-ssd when you need the IOPS and have measured it
pd-standard HDD, for throughput-oriented bulk data
hyperdisk newest, IOPS and throughput provisioned independently of size
Persistent disk performance scales with provisioned size — a 20 GB pd-balanced is slow because it is small, not because the type is wrong. That is the most common “the disk is slow” cause, and the fix is a bigger disk or Hyperdisk with explicit IOPS.
Zonal versus regional: a regional persistent disk replicates synchronously across two zones in the region and can be force-attached in the surviving zone. It costs roughly double and is the difference between a zonal VM and one that can be recovered elsewhere — the resilience lesson’s point, made concrete.
gcloud compute disks create data-1 --region=europe-west1 \
--replica-zones=europe-west1-b,europe-west1-c --size=200GB --type=pd-balanced
Resize is online and one-way: you can grow a disk (and then the filesystem) without downtime; you cannot shrink it.
Topic 3: Snapshots, Images and Machine Images
Three different artifacts that people conflate:
Snapshot — an incremental, regional backup of a disk. The first is full, later ones store only changed blocks, and each is independently restorable. This is your cross-zone escape hatch: a zonal disk plus a snapshot equals a disk in any zone.
gcloud compute resource-policies create snapshot-schedule daily-snap \
--region=europe-west1 --max-retention-days=14 \
--daily-schedule --start-time=02:00 --storage-location=eu
gcloud compute disks add-resource-policies data-1 --zone=europe-west1-b \
--resource-policies=daily-snap
Custom image — a bootable template for new VMs. This is what an image pipeline produces, and what a MIG’s instance template references.
Machine image — the whole VM: all disks, metadata, and configuration. Convenient for lift-and-shift and for cloning a snowflake you have not finished taming.
Snapshots are crash-consistent, not application-consistent. A snapshot of a running database is equivalent to pulling the power cord — usually recoverable, and not a backup strategy. Use the database’s own backup mechanism, or quiesce the filesystem for the instant of the snapshot.
Topic 4: Metadata, Startup Scripts and OS Login
The metadata server at 169.254.169.254 is how a VM learns about itself and gets credentials:
curl -H "Metadata-Flavor: Google" \
http://169.254.169.254/computeMetadata/v1/instance/service-accounts/default/token
Metadata-Flavor: Google is required, which is GCP’s equivalent of the IMDSv2 protection from the AWS module — a plain SSRF that only issues a bare GET cannot read it.
Startup scripts run as root on every boot, not just the first:
#!/bin/bash
set -euo pipefail
exec > >(tee /var/log/startup.log|logger -t startup -s 2>/dev/console) 2>&1
apt-get update && apt-get install -y nginx
gcloud secrets versions access latest --secret=app-config > /etc/app/config.json
systemctl enable --now nginx
Two things to know: output goes to the serial console and the log file you redirect it to, and a failing startup script does not fail the instance — the VM reaches RUNNING, joins the load balancer and serves errors. Anything essential must be reflected in a health check.
OS Login replaces SSH key metadata with IAM:
gcloud compute instances add-metadata app-1 --metadata enable-oslogin=TRUE
gcloud compute instances add-iam-policy-binding app-1 --zone=europe-west1-b \
--member=user:alice@acme.example --role=roles/compute.osLogin
Access is granted and revoked in IAM, sessions are audited, and there are no keys spread through instance metadata. Combined with IAP TCP forwarding, a VM needs no external IP and no open port 22:
gcloud compute ssh app-1 --zone=europe-west1-b --tunnel-through-iap
That pair is the GCP equivalent of AWS Session Manager, and it should be the default access path.
Topic 5: Availability Behaviour
Live migration moves a running VM to another host for maintenance with no reboot — the default, and one of the genuinely nice things about Compute Engine. Set --maintenance-policy=TERMINATE only when you must (GPUs, sole-tenant nodes with licensing constraints).
Automatic restart brings a VM back after a host failure, on by default.
Host maintenance events are visible in metadata, so a workload that needs to drain can react:
curl -H "Metadata-Flavor: Google" \
http://169.254.169.254/computeMetadata/v1/instance/maintenance-event
Spot VMs are the same machine at 60–91% off, reclaimable at any time with a 30-second notice delivered as a shutdown script and a metadata change. They belong behind a MIG, for workloads that tolerate interruption — which is the cost lesson’s point applied here.
Sole-tenant nodes give you a physical host to yourself, for licensing or compliance. Expensive, and occasionally the only way to satisfy a rule.
Topic 6: Operating a Fleet Rather Than a Pet
The end state is that no VM is configured by hand:
□ an image pipeline (Packer or Cloud Build) producing versioned custom images
□ instance templates referencing an image, not a family, for production
□ MIGs with auto-healing health checks — the MIG lesson covers this
□ startup scripts limited to last-mile configuration
□ Ops Agent installed in the image for metrics and logs
□ OS Login + IAP; no SSH keys in metadata, no external IPs
□ snapshot schedules on every disk holding state
□ labels on every instance for cost attribution
The Ops Agent is worth calling out: without it you get no memory, disk or process metrics, because the hypervisor cannot see inside the guest. “Memory usage is missing from Cloud Monitoring” is always this.
curl -sSO https://dl.google.com/cloudagents/add-google-cloud-ops-agent-repo.sh
sudo bash add-google-cloud-ops-agent-repo.sh --also-install
Try it yourself: create a VM with a local SSD, write a file to it, stop and start the instance, and look for the file. It is gone, with no warning at any point in the process — which is the single most surprising Compute Engine behaviour and the reason local SSD belongs to caches and scratch space only.
Common mistake: treating a VM you configured by hand as the source of truth, and taking a machine image of it as the “backup”. The image captures the state including the drift, nobody can reproduce how it got there, and the next person to need a second instance clones a snowflake. Build the image from a pipeline, keep the inputs in Git, and let the VM be disposable.