EC2 Instances: Lifecycle, Families and Placement

Reading an instance type name, what each state actually bills, why stop-start moves your host, and the bootstrapping choices that decide how fast a replacement enters service.

intermediate 22 min lesson hands-on task included

EC2 is the service everything else is built on, and the part that matters operationally is not “how do I launch an instance” but what the instance is: a virtual machine on a host you do not choose, with a lifecycle that has surprising properties at every transition.


Topic 1: The Lifecycle and What Each State Costs

running YOU PAY FOR instance-hours + EBS + EIP-in-use YOU KEEP everything stopped YOU PAY FOR EBS only (and unattached EIPs) YOU KEEP EBS, private IP, ENI terminated YOU PAY FOR nothing YOU KEEP only volumes with DeleteOnTermination=false stop terminate STOP → START MOVES THE HOST Instance store data is gone, the public IPv4 changes, and the instance may land on newer hardware. HIBERNATE IS THE EXCEPTION RAM is written to the encrypted root volume, so you pay for the bigger volume instead of the instance. THE COST MISTAKE PEOPLE ACTUALLY MAKE Stopping an instance for a quarter and paying for its 500 GB gp3 volume the whole time. Snapshot and delete instead.
Three states, three billing profiles, and one transition — stop then start — that quietly changes the host, the public address, and the contents of instance store.

Running — billed per second (60-second minimum) for the instance, plus EBS, plus any in-use Elastic IP.

Stopped — no instance charge. You still pay for the EBS volumes, and now you also pay for unattached Elastic IPs, which is AWS charging for the address you are holding without using. The private IP and the ENI survive; the public IPv4 does not.

Terminated — nothing. Volumes with DeleteOnTermination=true (the default for the root volume) are gone; others survive as orphans that keep billing until someone notices.

The stop-start transition is a migration. The instance is re-placed on a different physical host, which means:

  • Instance store (ephemeral NVMe) contents are gone. Not warned about, not recoverable.
  • The public IPv4 changes, unless it is an Elastic IP.
  • The instance may land on newer hardware, which is how “stop-start fixed it” happens after a degraded host.

That last point is also the fix for the most common EC2 support case: a failing system status check means AWS’s infrastructure has a problem with the host, and stop-start moves you off it. A failing instance status check means your OS is unhealthy — kernel panic, full disk, exhausted memory — and moving hosts will not help.

aws ec2 describe-instance-status --instance-ids i-0abc \
  --query 'InstanceStatuses[].[SystemStatus.Status,InstanceStatus.Status]'

Set up auto-recovery for the system-check case; it is a CloudWatch alarm on StatusCheckFailed_System with a recover action, and it costs nothing. Modern instance types have it on by default, which is worth verifying rather than assuming.

Hibernate is the exception to stop-start: RAM is written to the encrypted root volume and restored on start, so the process tree survives. You pay for the larger volume instead of the instance — useful for a workstation with a long warm-up, rarely useful for a server.


Topic 2: Reading an Instance Type

READING AN INSTANCE TYPE — m7gd.2xlarge m family family: m = balanced, c = compute, r = memory, i = storage IO, g/p = G 7 generation generation: higher is newer, usually cheaper per unit of work g processor processor: g = Graviton (ARM), a = AMD, i = Intel, blank = the default fo d extras extras: d = local NVMe, n = extra network, e = extra storage/memory 2xlarge size size: each step doubles vCPU, RAM, network and EBS bandwidth THE ONE LETTER WORTH ARGUING ABOUT Graviton is typically 20–40% cheaper per unit of work — if your image is multi-arch. SIZE IS NOT ONLY vCPU AND RAM Network and EBS bandwidth scale with size too, and "up to" figures are burst. A small instance can bottleneck on EBS while CPU sits at 15%.
Five fields in one string. The processor letter is the one worth arguing about in a design review — everything else follows from the workload.

m7gd.2xlarge decomposes as family m, generation 7, processor g, extra d, size 2xlarge.

Families, by what they optimise:

PrefixForNote
tBurstable general purposeCPU credits — see below
mBalancedThe default when you have no data
cCompute-optimisedHigher clock, less RAM per vCPU
r, x, zMemory-optimisedCaches, in-memory databases, JVMs
i, dStorage-optimisedLocal NVMe, very high IOPS
p, g, inf, trnAcceleratedGPU and purpose-built ML silicon

The t family deserves a warning. Burstable instances earn CPU credits when idle and spend them when busy. Run out and you are throttled to the baseline — 20% of a vCPU on a t3.medium — which presents as an application that becomes catastrophically slow while CPU utilisation reads a comfortable 20%. unlimited mode removes the cliff and adds a surcharge you will not notice until the bill. Alarm on CPUCreditBalance, or use a non-burstable family for anything that must be predictable.

The processor letter is the biggest lever. g is Graviton — AWS’s ARM silicon — and is typically 20–40% better price/performance for the same work. The requirement is that your entire image stack is multi-arch: base images, compiled dependencies, agents, sidecars. For interpreted and JVM workloads this is usually a one-line change to a build; for anything with native extensions it is a real migration. That difficulty inversion is why Graviton adoption is easy for a web tier and hard for a data-processing tier.

Size scales more than vCPU and RAM. Network bandwidth and EBS bandwidth scale with size too, and small sizes quote “up to” figures that are burst-only. A c6i.large starved of EBS throughput at 15% CPU is a very common and very confusing performance ticket.


Topic 3: Bootstrapping and Boot-to-Service Time

The metric that matters for an Auto Scaling group is not boot time; it is boot-to-service time — from RunInstances to serving traffic. Three approaches:

USER DATA ONLY          AMI is generic; a script installs everything at boot.
  boot 40s + install 3–8 min = slow scale-out, and it breaks when
  a package repository has a bad day.

FULLY BAKED AMI         Everything pre-installed and pre-configured.
  boot 40s + start 10s = fast, deterministic, and needs an AMI
  pipeline plus a new AMI for every change.

HYBRID (the usual answer)   AMI has the heavy dependencies; user data does
  the last mile of configuration.
  boot 40s + config 20s = fast enough, still flexible.

User data runs once, as root, at first boot. Two properties bite regularly:

  • Output goes to /var/log/cloud-init-output.log. That is the first place to look when an instance launches and does nothing.
  • A failing user-data script does not fail the launch. The instance reaches running, passes status checks, joins the target group, and serves errors. If a bootstrap step is essential, the readiness signal must come from the application, not from the instance state.
#!/bin/bash
set -euo pipefail
exec > >(tee /var/log/user-data.log|logger -t user-data -s 2>/dev/console) 2>&1

dnf install -y nginx
aws ssm get-parameter --name /app/prod/config --with-decryption \
  --query Parameter.Value --output text > /etc/app/config.json
systemctl enable --now nginx

Note where the configuration comes from: SSM Parameter Store, fetched at boot with the instance’s own role. Standard parameters are free, SecureString values are KMS-encrypted, and versions are retained. It means the AMI is identical across environments and the environment differences live in one place you can audit.

Build AMIs with a pipeline, not by hand. EC2 Image Builder or Packer, triggered on a schedule and on base-image updates, producing a versioned AMI whose ID goes into the launch template. A hand-built AMI is an undocumented artifact that nobody can reproduce in a year.


Topic 4: Placement Groups

By default AWS places instances wherever it likes within the AZ. Placement groups constrain that, in three shapes:

TypeBehaviourUse for
ClusterPacked onto the same high-bandwidth, low-latency rack segment, single AZHPC, tightly coupled workloads that need microsecond latency
SpreadEach instance on distinct hardware, up to 7 per AZA small number of critical peers — three etcd members, three Kafka brokers
PartitionGroups of instances on separate racks, up to 7 partitions per AZLarge distributed systems that are rack-aware: HDFS, Cassandra

Cluster placement maximises throughput and minimises resilience — one rack event takes everything. It also increases the chance of InsufficientInstanceCapacity, because AWS must find that much capacity in one place. Launch all instances in a single request to improve the odds.

Spread placement is the one most people should know about and few use. Three critical instances in a spread group cannot share a hardware failure domain. It is free, and the 7-per-AZ limit is the reason it does not scale to fleets.


Topic 5: Purchase Options as an Availability Decision

Pricing is covered properly in the Cloud Cost Optimization path. What belongs here is the reliability side of the same choice:

  • On-Demand — no commitment, no interruption. The baseline.
  • Spot — up to 90% off, with a two-minute interruption notice. Capacity is reclaimed when AWS wants it back.
  • Reserved Instances / Savings Plans — a billing construct, not a capacity guarantee. A Savings Plan does not reserve anything.
  • Capacity Reservations — the only option that actually reserves capacity in an AZ. Billed whether you use it or not, and the correct tool for a disaster-recovery region where “we will launch instances when we need them” must not fail.

Spot is production-viable when the workload can absorb interruption: stateless web tiers behind a load balancer, batch, CI runners, Kubernetes nodes running interruptible pods. The rules that make it work are diversification (many instance types across many AZs, so one pool’s reclamation is not your outage) and handling the notice — drain, deregister, checkpoint:

# Poll for the interruption notice from inside the instance
curl -s -H "X-aws-ec2-metadata-token: $TOKEN" \
  http://169.254.169.254/latest/meta-data/spot/instance-action
# 404 until it isn't; then {"action":"terminate","time":"2026-08-18T09:12:00Z"}

Two minutes is enough to deregister from a target group and finish in-flight requests. It is not enough to migrate state, which is the real constraint on where Spot belongs.


Topic 6: Launch Templates

Launch configurations are the deprecated predecessor; launch templates are what you should use, and the difference matters because ASGs, Spot fleets and EKS managed node groups all consume them.

A launch template is versioned, and an ASG references either a specific version or $Latest/$Default. Referencing a pinned version and updating deliberately is the safer pattern — $Latest means a template edit changes production at the next scale-out, hours later, with no deployment event to correlate against.

What belongs in the template, every time:

✓ AMI ID (from the pipeline, parameterised)
✓ Instance profile           — never keys
✓ Metadata options           — HttpTokens: required, hop limit 1
✓ Security groups            — referencing other SGs
✓ EBS mapping                — gp3, encrypted, sized deliberately
✓ Tags on instances AND volumes  — untagged volumes are how orphans hide
✓ User data                  — the last mile only
✓ Detailed monitoring        — 1-minute metrics, if you alarm on them

Tagging volumes through the template is worth calling out: tags on the instance do not propagate to its volumes automatically, and untagged volumes are exactly the ones that survive a termination and bill for a year.

Try it yourself: launch two instances from the same template, one with a user-data install and one from a baked AMI, and time both to first successful health check. The gap is what your Auto Scaling group pays every time it scales out — and during an incident, that gap is your recovery time.

Common mistake: treating stop as a cost optimisation for an idle instance with a large volume. A stopped instance with a 500 GB gp3 root volume still costs about $40 a month, forever, for nothing. Snapshot it, terminate it, and restore when needed — or admit it will never be needed and delete it.