Storage performance problems on AWS are almost always one of three things: the volume type is wrong, the instance cannot carry what the volume can deliver, or the workload is doing small random I/O against something optimised for sequential throughput. All three are measurable before they become incidents.
Topic 1: The Three Kinds of Storage
INSTANCE STORE NVMe physically attached to the host.
Fastest possible. Gone on stop, terminate, or host failure.
Included in the instance price. Cannot be detached.
EBS Network-attached block storage, replicated within one AZ.
Survives instance termination. Snapshot-able. Zonal.
One volume, one instance at a time (except io1/io2 Multi-Attach).
S3 / EFS / FSx Object and file storage. Not block devices; different lesson,
different failure modes.
Instance store is not a cheaper EBS. It is a different guarantee: enormous IOPS with no durability whatsoever. Correct uses are caches, scratch space, temporary shuffle data, and replicated datastores where the cluster provides the durability. Anything you would be upset to lose on a stop-start does not belong there.
EBS is network storage that looks like a disk. Every read and write crosses the network, which is why instance network characteristics constrain it, and why EBS-optimised instances (all current generations, by default) matter.
Topic 2: The Two Ceilings
Ceiling 1 — the volume:
| Type | Profile | Use for |
|---|---|---|
| gp3 | 3,000 IOPS and 125 MB/s baseline included; up to 16,000 IOPS and 1,000 MB/s, provisioned independently of size | The default for everything |
| gp2 (legacy) | 3 IOPS per GiB, burst bucket for small volumes | Migrate away |
| io2 Block Express | Up to 256,000 IOPS, sub-millisecond, 99.999% durability | Serious databases |
| st1 | Throughput-optimised HDD, sequential | Big log and data-lake volumes |
| sc1 | Cold HDD, cheapest | Archive you must mount |
gp3 over gp2, essentially always. gp2 ties performance to size, which forces you to buy 1 TB to get 3,000 IOPS. gp3 gives you 3,000 IOPS at any size and is around 20% cheaper per gigabyte. The change is online — modify-volume with no downtime:
aws ec2 modify-volume --volume-id vol-0abc --volume-type gp3
aws ec2 describe-volumes-modifications --volume-ids vol-0abc \
--query 'VolumesModifications[].[ModificationState,Progress]'
Do check the volumes where gp2’s size-derived IOPS exceeded 3,000 (anything above 1 TiB), because those need explicit provisioning on gp3 or they get slower.
Ceiling 2 — the instance. Every instance size has a maximum EBS bandwidth and IOPS, and smaller sizes only reach the headline number in bursts. If the instance caps at 6,000 IOPS, provisioning 16,000 on the volume buys nothing but the bill. Attaching more volumes does not help either — the ceiling is per instance, not per volume, which is the clean way to diagnose it.
Both ceilings appear in CloudWatch. VolumeQueueLength climbing while VolumeReadOps is flat means you are saturated; whether it is the volume or the instance is answered by comparing against the instance’s documented maximum. On Nitro instances, EBSIOBalance% and EBSByteBalance% show the instance-level burst bucket draining, which is the clearest possible signal that the instance, not the volume, is the constraint.
Topic 3: Snapshots — What They Actually Are
Snapshots are incremental and stored in S3 (in a service-managed bucket you never see). The first snapshot copies every used block; subsequent ones copy only changed blocks. Crucially, each snapshot is independently restorable — deleting an old snapshot never breaks a newer one, because blocks still referenced are retained. This is the opposite of an incremental-backup chain, and it means retention policy can be simple.
Three properties with operational consequences:
Restores are lazy. A volume created from a snapshot is immediately available, but blocks are fetched from S3 on first access. Early reads are slow — sometimes dramatically. Options: pre-warm by reading the whole device (fio --rw=read --bs=1M across it), or enable Fast Snapshot Restore, which removes the penalty and bills per snapshot per AZ per hour. FSR is expensive enough that it should be enabled for a specific recovery drill or a known launch event, not left on by default.
Snapshots are crash-consistent, not application-consistent. A snapshot of a running database is exactly equivalent to pulling the power cord: the filesystem journal will replay, and the database will run recovery. Usually it works. “Usually” is not a backup strategy. For a database, use the database’s own backup mechanism (RDS automated backups, pg_dump, an engine-level snapshot) or freeze the filesystem for the instant of the snapshot.
Snapshots are regional; copies are how you leave the region. copy-snapshot to another region is the primitive under most EBS-level DR. Copies can be re-encrypted with a different KMS key in the process, which is how you move data between accounts with separate key hierarchies.
Automate retention with Data Lifecycle Manager rather than a cron job on a bastion:
aws dlm create-lifecycle-policy \
--description "daily snapshots, 7-day retention" \
--state ENABLED --execution-role-arn arn:aws:iam::111122223333:role/dlm \
--policy-details file://policy.json
DLM is free, tag-based, and handles cross-region copy and cross-account share. A snapshot script on an instance is a thing that stops working silently when the instance is replaced.
Topic 4: Encryption
EBS encryption uses KMS with envelope encryption and costs no measurable performance on current instances. There is no reason not to enable it.
Turn on encryption by default at the account level, per region:
aws ec2 enable-ebs-encryption-by-default
aws ec2 modify-ebs-default-kms-key-id --kms-key-id alias/ebs-default
The rules that shape migrations:
- An unencrypted volume cannot be encrypted in place. The path is snapshot → copy the snapshot with encryption → create a volume from the copy → swap.
- A volume created from an encrypted snapshot is always encrypted.
- Sharing an encrypted snapshot cross-account requires a customer-managed key — the AWS-managed
aws/ebskey cannot be shared, and this is discovered at the least convenient moment. - The KMS key must exist in the same region as the volume. Cross-region copies re-encrypt with a key in the destination region.
Topic 5: Multi-Attach and What It Does Not Give You
io1 and io2 volumes support Multi-Attach: up to 16 instances in the same AZ attach the same volume simultaneously. What that does not provide is a shared filesystem. ext4 and XFS assume exclusive access; mounting one volume read-write from two instances corrupts it immediately and completely.
Multi-Attach is only for cluster-aware filesystems and applications that coordinate their own access — GFS2, OCFS2, or a database with its own clustering layer. If what you actually want is “several instances reading and writing the same files”, the answer is EFS (NFS, multi-AZ, elastic) or FSx (Windows or Lustre), not Multi-Attach.
Quick orientation on the file options, since they come up whenever block storage does not fit:
| EFS | FSx for Windows | FSx for Lustre | |
|---|---|---|---|
| Protocol | NFSv4 | SMB | Lustre |
| Scope | Multi-AZ regional | Single or multi-AZ | Single AZ |
| Scale | Petabytes, elastic | TBs | Petabytes, very high throughput |
| For | Shared app data, home dirs | AD-integrated Windows workloads | HPC, ML training against S3 |
EFS’s own trap is worth one line: the default Bursting throughput mode ties throughput to stored size, so a small EFS filesystem under sustained load exhausts burst credits and becomes slow in a way that looks like a network problem. Elastic throughput mode removes that failure mode and bills for what you use.
Topic 6: Finding the Waste and the Risk
Storage accumulates quietly. Four one-line audits worth running monthly:
# Unattached volumes — paying for nothing
aws ec2 describe-volumes --filters Name=status,Values=available \
--query 'Volumes[].[VolumeId,Size,CreateTime]' --output table
# Still on gp2 — a free ~20% saving and usually better performance
aws ec2 describe-volumes --filters Name=volume-type,Values=gp2 \
--query 'Volumes[].[VolumeId,Size,Attachments[0].InstanceId]' --output table
# Unencrypted volumes
aws ec2 describe-volumes --filters Name=encrypted,Values=false \
--query 'Volumes[].[VolumeId,Attachments[0].InstanceId]' --output table
# Snapshots older than your retention policy claims to allow
aws ec2 describe-snapshots --owner-ids self \
--query 'Snapshots[?StartTime<`2026-01-01`].[SnapshotId,VolumeSize,StartTime]' \
--output table
The first query is the one that pays for the exercise, every time. Volumes detached during an incident six months ago, still billing, still holding data nobody has classified.
Try it yourself: create a 100 GiB gp2 and a 100 GiB gp3 and fio them both. gp2 delivers 300 IOPS at that size, gp3 delivers 3,000, and gp3 costs less. That single measurement is the whole argument for the migration.
Common mistake: treating a snapshot as a backup without ever restoring one. A snapshot you have never restored is a hypothesis about a file format. Restore one on a schedule, mount it, and check the data is what you expect — that exercise is also the only honest way to know your actual recovery time, lazy-load penalty included.