Hardening Containers

Non-root users, dropped capabilities, read-only root filesystems, secrets that never reach a layer, and the two mounts that hand an attacker your host.

advanced 22 min lesson hands-on task included

Container isolation is a kernel-enforced boundary, not a hardware one (lesson 4). Every syscall a container makes is handled by your kernel. The defaults are reasonably safe; the defaults are also not what you want in production, and the gap between them is about eight flags.


Topic 1: Do Not Run as Root

By default the process inside a container runs as UID 0, and without user-namespace remapping that is the same UID 0 as on the host. It is restricted by dropped capabilities and a seccomp profile, so it is not full host root — but it is one kernel bug or one careless mount away from being exactly that.

Fix it in the image:

FROM node:22-alpine
WORKDIR /app
COPY --chown=node:node . .
RUN npm ci --omit=dev
USER node                        # many official images ship a ready-made non-root user
CMD ["node", "server.js"]

Or create one explicitly, with a high UID so it cannot collide with a meaningful host account:

RUN addgroup -g 10001 -S app && adduser -u 10001 -S -G app app
USER 10001:10001

Prefer the numeric form in USER. A username has to be resolved against /etc/passwd inside the image, which does not exist in distroless or scratch images, and Kubernetes’ runAsNonRoot check cannot verify a name.

And enforce it at run time as well, because you will inherit images you did not build:

docker run -d --user 10001:10001 myimage

The two things that break when you stop being root, and their fixes:

  • Binding ports below 1024. Listen on 8080 and publish it as -p 80:8080. There is never a reason for the process itself to bind 80.
  • Writing to paths owned by root. COPY --chown at build time, and mount a volume or tmpfs for anything written at run time.

Topic 2: Drop Capabilities

Linux splits root’s powers into ~40 capabilities. Docker already drops most of them, but it leaves a default set that almost nothing needs. Drop everything and add back only what is proven necessary:

docker run -d \
  --cap-drop ALL \
  --cap-add NET_BIND_SERVICE \
  --security-opt no-new-privileges \
  myimage
CapabilityGrantsUsually needed?
NET_BIND_SERVICEBind ports below 1024Only if you refuse to use a high port
CHOWN, FOWNER, DAC_OVERRIDEBypass file ownership and permission checksAlmost never at run time
SETUID, SETGIDChange process UID/GIDOnly if the app drops privileges itself
NET_RAWRaw sockets — this is what lets a container pingNo. Drop it
SYS_ADMINVery close to full rootNever. If something asks, redesign

--security-opt no-new-privileges blocks setuid binaries from elevating privileges. It costs nothing and closes a whole class of escalation. Set it on everything.

The blunt instrument to avoid entirely:

docker run --privileged ...       # disables capabilities, device restrictions and seccomp

--privileged is “run this as root on the host with extra steps”. It shows up in tutorials for Docker-in-Docker and hardware access, and there is nearly always a narrower answer — --device for one device, a specific --cap-add for one capability.


Topic 3: Read-Only Root Filesystem

If your application does not need to write to its own filesystem, do not let it:

docker run -d \
  --read-only \
  --tmpfs /tmp:rw,noexec,nosuid,size=64m \
  --tmpfs /run:rw,noexec,nosuid,size=16m \
  -v applogs:/var/log/app \
  myimage

This defeats the most common post-exploitation step: dropping a tool or a web shell onto disk. The noexec on the tmpfs mounts closes the obvious workaround.

Making it work is an exercise in finding out what your application actually writes. Start with --read-only, watch it fail, and add a tmpfs or a volume for each path — /tmp, a PID file directory, a cache directory. Compose expresses the same thing:

    read_only: true
    tmpfs:
      - /tmp:size=64m,noexec,nosuid

Two related controls in the same family:

docker run --security-opt seccomp=/path/to/profile.json ...   # restrict syscalls
docker run --security-opt apparmor=my-profile ...             # MAC, on Debian/Ubuntu

Docker applies a default seccomp profile that already blocks around 44 dangerous syscalls. Do not disable it (--security-opt seccomp=unconfined) to make something work without understanding exactly which syscall it needs.


Topic 4: Secrets

State the rule first: anything in an image is readable by anyone with the image.

The wrong ways, and why each fails:

ARG API_TOKEN                              # visible in `docker history`
ENV DB_PASSWORD=hunter2                    # visible in `docker inspect`, and in the env of every process
COPY id_rsa /root/.ssh/id_rsa
RUN git clone ... && rm /root/.ssh/id_rsa  # still recoverable from the earlier layer (lesson 5)

That third one is the nastiest, because it looks handled. Copy-on-write means the deletion is a whiteout in an upper layer; the key bytes are still in the layer below, and docker save plus tar extracts them in about thirty seconds.

The right ways:

At build time — BuildKit secrets, mounted for one instruction, never written to a layer:

# syntax=docker/dockerfile:1
RUN --mount=type=secret,id=npmrc,target=/root/.npmrc npm ci
docker build --secret id=npmrc,src=$HOME/.npmrc -t app:1.0 .

At run time — a file, not an environment variable:

docker run -d --mount type=bind,source=/etc/app/creds,target=/run/secrets,readonly myimage

Environment variables leak in more places than people expect: docker inspect, /proc/<pid>/environ for anything else in the namespace, crash dumps, error reporters that helpfully serialise the environment, and child processes that inherit them. Files can be permission-restricted and re-read after rotation. Prefer files, and prefer a real secret manager (Vault, AWS Secrets Manager, SOPS-encrypted files) over both.

Check what you already shipped:

docker history --no-trunc myimage:1.0 | grep -iE 'password|token|key|secret'
docker image inspect myimage:1.0 --format '{{json .Config.Env}}'

If you find something, rotating the credential is the fix. Deleting the image is not — it has been pulled.


Topic 5: The Two Mounts That Are Game Over

/var/run/docker.sock. Mounting the Docker socket into a container gives that container control of the daemon. The daemon can start a new container that bind-mounts the host’s / — no exploit, no vulnerability, just the API working as designed. It appears constantly in CI tutorials and monitoring agent docs.

docker run -v /var/run/docker.sock:/var/run/docker.sock ...   # this is root on the host

If a container genuinely needs to build images, use a rootless builder (BuildKit’s buildkitd, or Kaniko) rather than the host socket. If it needs to observe containers, most monitoring agents can read from a read-only socket proxy that allows only the endpoints they need.

The host filesystem. -v /:/host is the same outcome, stated plainly. Watch for the softer versions too: -v /etc:/etc:ro still exposes every credential in /etc, and -v /proc:/host/proc exposes far more than most people assume.


Topic 6: Image Provenance and Scanning

Where did the base image come from? Official images in Docker Hub’s library namespace are maintained by Docker and upstream projects. Everything else is maintained by whoever pushed it. FROM randomuser/java-app:latest in a production Dockerfile is trusting a stranger with your production runtime, on a tag that can change under you.

Pin by digest where it matters:

FROM node:22-alpine@sha256:1a2b3c4d...

Scan continuously, not once. A clean image today acquires CVEs as the world discovers them in libraries you already shipped:

docker scout cves acme/api:1.4.2
trivy image --severity HIGH,CRITICAL acme/api:1.4.2
grype acme/api:1.4.2

Wire this into CI with a failure threshold. And note the highest-leverage remediation is usually not patching — it is shrinking the base image. Most CVEs in a typical application image are in operating-system packages the application never calls. Moving from ubuntu to -slim to distroless routinely takes a finding count from hundreds to single digits, which is lesson 7 paying a second dividend.

Rootless Docker is worth knowing about as the structural version of Topic 1: the daemon itself runs as an unprivileged user, using user namespaces so container root maps to your UID on the host. Some features are limited (port binding below 1024, some networking and storage drivers), but for development machines and many CI runners it removes an entire class of risk.

Try it yourself: Build a trivially “leaky” image — ARG TOKEN=secret123 plus a RUN echo $TOKEN > /tmp/t && rm /tmp/t. Then run docker history --no-trunc and find the token in plain text, and docker save it, extract the tar, and find the deleted file in an earlier layer. Doing this once makes the rule permanent.

Common mistake: Hardening the run-time flags and leaving the image running as root. --cap-drop ALL on a root process is much weaker than a non-root process with default capabilities. The user is the control that matters most; the rest is defence in depth on top of it.