Images, Layers and Copy-on-Write

Why an image is a stack of read-only layers, what the storage driver unions together to make one filesystem, and how the thin writable layer decides what survives a container's death.

intermediate 20 min lesson hands-on task included

An image is not a tarball of a filesystem. It is an ordered stack of read-only layers, and almost every practical fact about Docker — build speed, push size, disk usage, why your data disappeared — falls out of that one design decision.


Topic 1: Layers Are Created by Instructions

Each instruction in a Dockerfile that changes the filesystem creates a new layer. RUN, COPY and ADD create layers. Metadata-only instructions — ENV, LABEL, WORKDIR, EXPOSE, CMD, ENTRYPOINT — do not add filesystem layers; they change the image configuration.

FROM alpine                                    # base layers, inherited
RUN mkdir /mycontent && touch /mycontent/1.txt # one new layer
docker build -t layerdemo .
docker history layerdemo          # every layer, its size, and the instruction that made it
docker image inspect layerdemo --format '{{len .RootFS.Layers}}'

docker history is the tool. It shows you, top to bottom, which instruction produced which layer and how many bytes it added. When an image is unexpectedly 1.4 GB, docker history names the line that did it, usually within five seconds.

Compare docker image inspect alpine and docker image inspect layerdemo and you will see the layer lists are identical except for the one extra entry. The base layers are not copied — they are shared.


Topic 2: Sharing Is Why Everything Is Fast

Layers are content-addressed and immutable, so identical layers are stored once and reused everywhere.

If image X is 500 MB and you build a second image on the same base, the second image does not cost another 500 MB. It costs the size of what you added. This is why:

  • A rebuild after a code change pushes kilobytes, not gigabytes.
  • Twelve microservices on the same node:22-alpine base share one copy of that base on every host.
  • Pulling the fifth image from a registry is far faster than pulling the first.
docker system df                  # images / containers / volumes / build cache, with reclaimable sizes
docker system df -v               # per-image and per-container detail

docker system df is one of the most under-used commands in the toolset. The RECLAIMABLE column tells you exactly how much of your disk is dead weight before you prune anything.

The build cache follows the same rule:

When you rebuild, Docker walks your Dockerfile instruction by instruction. For each one it asks: is there a cached layer produced by this exact instruction on top of this exact parent? If yes, reuse it. The first miss invalidates every layer after it.

That is the entire reason for Dockerfile ordering discipline. Put the things that change rarely (base image, system packages, dependency manifests) near the top, and the things that change every commit (your source code) near the bottom. Lesson 6 turns this into concrete rules.


Topic 3: The Thin Writable Layer

Image layers are read-only. Always. When you create a container, Docker adds one more layer on top — a thin read-write layer — and that container’s writes go there.

┌──────────────────────────────┐
│  container R/W layer  ← writes go here, deleted with the container
├──────────────────────────────┤
│  layer: COPY app/            │
│  layer: RUN apt-get install  │  read-only, shared between every
│  layer: base image           │  container started from this image
└──────────────────────────────┘

Ten containers from one image means one copy of the read-only layers and ten small writable layers.

docker diff shows you exactly what a container has changed relative to its image:

docker diff api
# C /etc
# A /etc/nginx/conf.d/custom.conf     A = added
# C /var/log/nginx/access.log         C = changed
# D /tmp/scratch                      D = deleted

When the container is deleted, that layer is deleted with it. This is the mechanism behind the warning in lesson 1: everything you install by hand inside a running container is gone the moment the container is removed. It is also why lesson 8 exists.


Topic 4: Copy-on-Write

If the layers below are read-only, how does editing an existing file work?

docker run --name one -it layerdemo /bin/sh
# inside:
cat /mycontent/1.txt
echo "hello" >> /mycontent/1.txt

The file /mycontent/1.txt lives in a read-only layer. On first write, the storage driver copies the whole file up into the writable layer and applies the change there. Every subsequent read of that path resolves to the copy. This is copy-on-write.

Three behaviours follow, and all three bite in production:

First write to a large file is slow. Copying a 2 GB file up before modifying one byte costs a 2 GB copy. Databases and anything doing heavy random writes to image-resident files should be on a volume, not on the writable layer — performance is the reason as much as persistence.

Deleting a file does not shrink the image. Deletion in an upper layer records a whiteout marker; the bytes in the lower layer are still there. This is why RUN apt-get install ... && RUN rm -rf /var/lib/apt/lists/* as two separate instructions saves nothing — the files exist in the layer below the delete. Combine them into one RUN and the files are never committed in the first place.

Secrets deleted in a later layer are still recoverable. Same mechanism, much worse consequence. If you COPY a private key and RUN rm it in a later instruction, anyone with the image can extract it from the earlier layer. Lesson 12 covers what to do instead.


Topic 5: Storage Drivers

Something has to present a stack of layers to the container as one ordinary filesystem rooted at /. That something is the storage driver, and it does so with a union mount.

docker info | grep -i 'storage driver'
DriverStatus
overlay2The default and the right answer on every modern Linux distribution
aufsThe pre-18.06 default. Historical interest only
devicemapperBlock-level, for old RHEL/CentOS. Slow and operationally awkward
btrfs / zfsAvailable if your host filesystem is one of those, with snapshot benefits

Everything Docker stores lives under /var/lib/docker:

sudo ls /var/lib/docker             # containers, image, volumes, overlay2, network
sudo du -sh /var/lib/docker/*       # where the disk actually went

/var/lib/docker/overlay2 is the image and container layer store. Never edit anything under /var/lib/docker by hand — the daemon’s metadata will disagree with the filesystem and you will get errors that make no sense. Use docker commands or, in a genuine emergency, stop the daemon first.

Inside the container, df -h will show the root filesystem mounted as overlay. That is the union of every layer, presented as one disk.


Topic 6: Tags, Digests and Cleaning Up

An image reference is repository:tag, and the tag is mutable. nginx:1.27 can point at different bytes next month. A digest is immutable — it is the content hash:

docker images --digests
docker pull nginx@sha256:c5a1f5a3...        # exactly these bytes, forever

Use tags for humans and digests where you need certainty (production deploys, compliance, reproducible builds). And never use :latest for anything real: it is a tag like any other, it moves without warning, and it makes rollback meaningless because there is nothing to roll back to.

Now the cleanup commands, in ascending order of danger:

docker image prune                # dangling images only (untagged, no children) — safe
docker image prune -a             # every image not used by a container — aggressive
docker container prune            # all stopped containers
docker builder prune              # the build cache, which grows without limit on CI hosts
docker system prune               # stopped containers + unused networks + dangling images + build cache
docker system prune -a --volumes  # also removes unused images AND VOLUMES — read that twice

docker system prune -a --volumes deletes data. Volumes not currently attached to a container are removed, and that includes the database volume whose container you stopped ten minutes ago. On a CI runner it is routine housekeeping. On a host with anything stateful, run docker volume ls first and know what you are about to lose.

Try it yourself: Build an image, then add a RUN mkdir /hello line and rebuild. Watch which steps say CACHED and which re-run. Then move the same line to the top of the Dockerfile, rebuild, and watch the entire cache below it invalidate. That is the whole cost of bad instruction ordering, demonstrated in one edit.

Common mistake: Trying to shrink an image by adding a RUN rm -rf instruction at the end. It never works — copy-on-write whiteouts hide the files without removing the bytes, and the image gets larger by one layer. Delete in the same RUN that created the files, or use a multi-stage build (lesson 7) so the files never reach the final image at all.