EC2 is the service everything else is built on, and the part that matters operationally is not “how do I launch an instance” but what the instance is: a virtual machine on a host you do not choose, with a lifecycle that has surprising properties at every transition.
Topic 1: The Lifecycle and What Each State Costs
Running — billed per second (60-second minimum) for the instance, plus EBS, plus any in-use Elastic IP.
Stopped — no instance charge. You still pay for the EBS volumes, and now you also pay for unattached Elastic IPs, which is AWS charging for the address you are holding without using. The private IP and the ENI survive; the public IPv4 does not.
Terminated — nothing. Volumes with DeleteOnTermination=true (the default for the root volume) are gone; others survive as orphans that keep billing until someone notices.
The stop-start transition is a migration. The instance is re-placed on a different physical host, which means:
- Instance store (ephemeral NVMe) contents are gone. Not warned about, not recoverable.
- The public IPv4 changes, unless it is an Elastic IP.
- The instance may land on newer hardware, which is how “stop-start fixed it” happens after a degraded host.
That last point is also the fix for the most common EC2 support case: a failing system status check means AWS’s infrastructure has a problem with the host, and stop-start moves you off it. A failing instance status check means your OS is unhealthy — kernel panic, full disk, exhausted memory — and moving hosts will not help.
aws ec2 describe-instance-status --instance-ids i-0abc \
--query 'InstanceStatuses[].[SystemStatus.Status,InstanceStatus.Status]'
Set up auto-recovery for the system-check case; it is a CloudWatch alarm on StatusCheckFailed_System with a recover action, and it costs nothing. Modern instance types have it on by default, which is worth verifying rather than assuming.
Hibernate is the exception to stop-start: RAM is written to the encrypted root volume and restored on start, so the process tree survives. You pay for the larger volume instead of the instance — useful for a workstation with a long warm-up, rarely useful for a server.
Topic 2: Reading an Instance Type
m7gd.2xlarge decomposes as family m, generation 7, processor g, extra d, size 2xlarge.
Families, by what they optimise:
| Prefix | For | Note |
|---|---|---|
t | Burstable general purpose | CPU credits — see below |
m | Balanced | The default when you have no data |
c | Compute-optimised | Higher clock, less RAM per vCPU |
r, x, z | Memory-optimised | Caches, in-memory databases, JVMs |
i, d | Storage-optimised | Local NVMe, very high IOPS |
p, g, inf, trn | Accelerated | GPU and purpose-built ML silicon |
The t family deserves a warning. Burstable instances earn CPU credits when idle and spend them when busy. Run out and you are throttled to the baseline — 20% of a vCPU on a t3.medium — which presents as an application that becomes catastrophically slow while CPU utilisation reads a comfortable 20%. unlimited mode removes the cliff and adds a surcharge you will not notice until the bill. Alarm on CPUCreditBalance, or use a non-burstable family for anything that must be predictable.
The processor letter is the biggest lever. g is Graviton — AWS’s ARM silicon — and is typically 20–40% better price/performance for the same work. The requirement is that your entire image stack is multi-arch: base images, compiled dependencies, agents, sidecars. For interpreted and JVM workloads this is usually a one-line change to a build; for anything with native extensions it is a real migration. That difficulty inversion is why Graviton adoption is easy for a web tier and hard for a data-processing tier.
Size scales more than vCPU and RAM. Network bandwidth and EBS bandwidth scale with size too, and small sizes quote “up to” figures that are burst-only. A c6i.large starved of EBS throughput at 15% CPU is a very common and very confusing performance ticket.
Topic 3: Bootstrapping and Boot-to-Service Time
The metric that matters for an Auto Scaling group is not boot time; it is boot-to-service time — from RunInstances to serving traffic. Three approaches:
USER DATA ONLY AMI is generic; a script installs everything at boot.
boot 40s + install 3–8 min = slow scale-out, and it breaks when
a package repository has a bad day.
FULLY BAKED AMI Everything pre-installed and pre-configured.
boot 40s + start 10s = fast, deterministic, and needs an AMI
pipeline plus a new AMI for every change.
HYBRID (the usual answer) AMI has the heavy dependencies; user data does
the last mile of configuration.
boot 40s + config 20s = fast enough, still flexible.
User data runs once, as root, at first boot. Two properties bite regularly:
- Output goes to
/var/log/cloud-init-output.log. That is the first place to look when an instance launches and does nothing. - A failing user-data script does not fail the launch. The instance reaches
running, passes status checks, joins the target group, and serves errors. If a bootstrap step is essential, the readiness signal must come from the application, not from the instance state.
#!/bin/bash
set -euo pipefail
exec > >(tee /var/log/user-data.log|logger -t user-data -s 2>/dev/console) 2>&1
dnf install -y nginx
aws ssm get-parameter --name /app/prod/config --with-decryption \
--query Parameter.Value --output text > /etc/app/config.json
systemctl enable --now nginx
Note where the configuration comes from: SSM Parameter Store, fetched at boot with the instance’s own role. Standard parameters are free, SecureString values are KMS-encrypted, and versions are retained. It means the AMI is identical across environments and the environment differences live in one place you can audit.
Build AMIs with a pipeline, not by hand. EC2 Image Builder or Packer, triggered on a schedule and on base-image updates, producing a versioned AMI whose ID goes into the launch template. A hand-built AMI is an undocumented artifact that nobody can reproduce in a year.
Topic 4: Placement Groups
By default AWS places instances wherever it likes within the AZ. Placement groups constrain that, in three shapes:
| Type | Behaviour | Use for |
|---|---|---|
| Cluster | Packed onto the same high-bandwidth, low-latency rack segment, single AZ | HPC, tightly coupled workloads that need microsecond latency |
| Spread | Each instance on distinct hardware, up to 7 per AZ | A small number of critical peers — three etcd members, three Kafka brokers |
| Partition | Groups of instances on separate racks, up to 7 partitions per AZ | Large distributed systems that are rack-aware: HDFS, Cassandra |
Cluster placement maximises throughput and minimises resilience — one rack event takes everything. It also increases the chance of InsufficientInstanceCapacity, because AWS must find that much capacity in one place. Launch all instances in a single request to improve the odds.
Spread placement is the one most people should know about and few use. Three critical instances in a spread group cannot share a hardware failure domain. It is free, and the 7-per-AZ limit is the reason it does not scale to fleets.
Topic 5: Purchase Options as an Availability Decision
Pricing is covered properly in the Cloud Cost Optimization path. What belongs here is the reliability side of the same choice:
- On-Demand — no commitment, no interruption. The baseline.
- Spot — up to 90% off, with a two-minute interruption notice. Capacity is reclaimed when AWS wants it back.
- Reserved Instances / Savings Plans — a billing construct, not a capacity guarantee. A Savings Plan does not reserve anything.
- Capacity Reservations — the only option that actually reserves capacity in an AZ. Billed whether you use it or not, and the correct tool for a disaster-recovery region where “we will launch instances when we need them” must not fail.
Spot is production-viable when the workload can absorb interruption: stateless web tiers behind a load balancer, batch, CI runners, Kubernetes nodes running interruptible pods. The rules that make it work are diversification (many instance types across many AZs, so one pool’s reclamation is not your outage) and handling the notice — drain, deregister, checkpoint:
# Poll for the interruption notice from inside the instance
curl -s -H "X-aws-ec2-metadata-token: $TOKEN" \
http://169.254.169.254/latest/meta-data/spot/instance-action
# 404 until it isn't; then {"action":"terminate","time":"2026-08-18T09:12:00Z"}
Two minutes is enough to deregister from a target group and finish in-flight requests. It is not enough to migrate state, which is the real constraint on where Spot belongs.
Topic 6: Launch Templates
Launch configurations are the deprecated predecessor; launch templates are what you should use, and the difference matters because ASGs, Spot fleets and EKS managed node groups all consume them.
A launch template is versioned, and an ASG references either a specific version or $Latest/$Default. Referencing a pinned version and updating deliberately is the safer pattern — $Latest means a template edit changes production at the next scale-out, hours later, with no deployment event to correlate against.
What belongs in the template, every time:
✓ AMI ID (from the pipeline, parameterised)
✓ Instance profile — never keys
✓ Metadata options — HttpTokens: required, hop limit 1
✓ Security groups — referencing other SGs
✓ EBS mapping — gp3, encrypted, sized deliberately
✓ Tags on instances AND volumes — untagged volumes are how orphans hide
✓ User data — the last mile only
✓ Detailed monitoring — 1-minute metrics, if you alarm on them
Tagging volumes through the template is worth calling out: tags on the instance do not propagate to its volumes automatically, and untagged volumes are exactly the ones that survive a termination and bill for a year.
Try it yourself: launch two instances from the same template, one with a user-data install and one from a baked AMI, and time both to first successful health check. The gap is what your Auto Scaling group pays every time it scales out — and during an incident, that gap is your recovery time.
Common mistake: treating stop as a cost optimisation for an idle instance with a large volume. A stopped instance with a 500 GB gp3 root volume still costs about $40 a month, forever, for nothing. Snapshot it, terminate it, and restore when needed — or admit it will never be needed and delete it.