Debugging Containers Under Pressure

A repeatable triage order for a container that will not start, will not stay up, or cannot be reached — plus the exit codes worth memorising and the disk problem that causes half of them.

advanced 22 min lesson hands-on task included

The commands in this lesson are all ones you have already met. What is new is the order, because the fastest way to waste an hour is to start with the interesting hypothesis instead of the cheap one.


Topic 1: The Triage Order

Always the same four questions, in this order:

docker ps -a              # 1. Does it exist, and what state is it in?
docker logs <name>        # 2. What did it say?
docker inspect <name>     # 3. What was it actually configured with?
docker exec -it <name> sh # 4. What does the world look like from inside?

Step 1 costs nothing and eliminates half the possibilities. A container in exited is a different problem from one in restarting, which is different again from one that is running but unreachable.

Step 2 is where the answer usually is. It is astonishing how often “the container is broken” means the application printed a clear error message that nobody read.

Step 3 catches the configuration mistakes — the wrong environment variable, a mount that did not apply, a limit that is lower than you meant. docker inspect is verbose, so query it:

docker inspect api --format '{{.State.Status}} {{.State.ExitCode}} {{.State.OOMKilled}}'
docker inspect api --format '{{json .Config.Env}}' | tr ',' '\n'
docker inspect api --format '{{json .Mounts}}'
docker inspect api --format '{{json .NetworkSettings.Networks}}'
docker inspect api | grep -i memory

Step 4 only after the first three, because it is the slowest and it changes the container’s state.


Topic 2: Exit Codes Worth Knowing

docker ps -a shows the exit code in its STATUS column, and a handful of them are diagnostic on sight:

CodeMeaningFirst move
0Clean exit — the process finishedYour command is not a long-running server (lesson 3)
1Generic application errorRead the logs
125The daemon failed — bad flag, bad optionRe-read your docker run line
126Command found but not executableCheck the file’s permission bit
127Command not foundWrong path, or the binary is not in this base image
137128 + 9 (SIGKILL)Almost always OOM. Check .State.OOMKilled
139128 + 11 (SIGSEGV)Segfault — often an architecture mismatch
143128 + 15 (SIGTERM)Orderly shutdown; something asked it to stop

The 125/126/127 trio is worth internalising because it separates your problem from Docker’s. 125 means Docker refused; 126 and 127 mean Docker started fine and your command was wrong.

docker inspect api --format '{{.State.ExitCode}} OOMKilled={{.State.OOMKilled}} Error={{.State.Error}}'

For 137 specifically, confirm rather than assume. OOMKilled: true means the cgroup limit killed it. OOMKilled: false with 137 means something else sent SIGKILL — an orchestrator, an operator, or a failed graceful stop that timed out.


Topic 3: Container Starts Then Immediately Exits

The most common failure, and the rule from lesson 3 covers most of it: the container lives as long as PID 1 lives.

Walk it:

docker ps -a                         # exit code?
docker logs --tail 50 <name>         # last words
docker inspect <name> --format '{{json .Config.Cmd}} {{json .Config.Entrypoint}}'

The four causes, in frequency order:

  1. The command completes. docker run alpine echo hi is not broken. Neither is a Dockerfile whose CMD runs a script that finishes.
  2. The process daemonises. nginx instead of nginx -g 'daemon off;'; httpd instead of httpd -DFOREGROUND. PID 1 forks a background process and exits, taking the container with it.
  3. A trailing argument replaced CMD. docker run -d myapp echo hi runs the echo instead of the app.
  4. A missing file or config. Exit 1, and the logs will say so.

If the logs are empty, the process died before writing anything — usually cause 3 or a missing binary. Override the entrypoint to get a shell and look around:

docker run --rm -it --entrypoint sh myimage
# then run the real command by hand and watch it fail

That trick — replacing the entrypoint with a shell — is the single most useful debugging move for a container that will not start.


Topic 4: Restart Loops and OOM

A container cycling through restarting is a restart policy fighting a process that keeps dying. The restart is not the bug; the exit is.

docker ps -a                                                 # RESTARTS count and STATUS
docker logs --tail 100 <name>                                # current attempt
docker inspect <name> --format '{{.RestartCount}} {{.State.ExitCode}}'
docker events --filter container=<name> --since 10m          # the pattern over time

Note that unlike Kubernetes, Docker keeps only the current container’s logs — there is no --previous. If the restarts are fast, get the logs to a driver that persists them before you lose the evidence, or set the restart policy to no temporarily so the failed container sits still while you inspect it.

For memory specifically:

docker stats --no-stream                                     # current usage vs limit
docker inspect <name> --format '{{.HostConfig.Memory}} {{.State.OOMKilled}}'
dmesg -T | grep -i 'killed process'                          # the kernel's own record

dmesg is the authoritative source. The kernel logs every OOM kill with the process name and how much it was using, on the host — which is exactly the evidence you need when the application logs stop mid-sentence with no error.

And recall the trap from lesson 4: a runtime that reads /proc/meminfo sees the host’s memory, not its cgroup limit. If a JVM or Node process is being killed well below what it thinks it has, set its own heap limit to match the container limit rather than raising the container limit forever.


Topic 5: It Is Running But I Cannot Reach It

Work outward from the process, not inward from the browser. Each step eliminates a layer.

# 1. Is the process actually listening, and on which address?
docker exec api sh -c 'netstat -tlnp 2>/dev/null || ss -tlnp'

# 2. Does it answer from INSIDE its own namespace?
docker exec api wget -qO- http://localhost:8080/health

# 3. Does it answer from another container on the same network?
docker run --rm --network mynet curlimages/curl -s http://api:8080/health

# 4. Is the port actually published, and where?
docker port api
docker ps                                # PORTS column

# 5. Does it answer from the host?
curl -v http://localhost:8080/health

The step that fails names the layer:

  • Fails at 1 — the application is not listening. Read its logs and config.
  • Listening on 127.0.0.1 instead of 0.0.0.0 — the single most common networking bug in containers. A process bound to loopback inside its own namespace is unreachable from outside it, no matter what you publish. Bind to 0.0.0.0.
  • Fails at 3 — a network problem. Are both containers on the same user-defined network? Is the name spelled right? Remember the default bridge has no DNS (lesson 9).
  • Fails at 4 — you did not publish the port, or you published the wrong side of the colon.
  • Fails at 5 only — host firewall, or you published to 127.0.0.1 and are testing from elsewhere.

For anything deeper, join the target’s network namespace with a container that actually has tools, rather than installing tcpdump into production:

docker run --rm -it --network container:api nicolaka/netshoot
# now dig, tcpdump, ss, curl and iperf all see exactly what `api` sees

DNS problems specifically:

docker exec api cat /etc/resolv.conf      # should show 127.0.0.11 on a user-defined network
docker exec api nslookup db
docker network inspect mynet              # who is actually attached

Topic 6: Build Failures

Build errors are their own category because the failing thing no longer exists when you see the error.

docker build --progress=plain -t app:dev .        # full output, no collapsing
docker build --no-cache -t app:dev .              # rule out a stale cached layer

--progress=plain is the flag to reach for first. BuildKit’s default output hides the command output you need.

The trick for a mid-build failure is that every successful step left an image behind. Build up to the last good stage and open a shell in it:

docker build --target builder -t app:build .
docker run --rm -it app:build sh                  # inspect the state the failing step saw

The recurring causes:

  • COPY fails with “no such file or directory” — the path is relative to the build context, not to the Dockerfile, and .dockerignore may be excluding it. Check both.
  • The context is enormous — “Sending build context to Docker daemon 3.4GB” means you have no .dockerignore (lesson 6).
  • A step that used to work now fails — an unpinned base image or unpinned packages moved underneath you. This is the argument for pinning, arriving late.
  • no space left on device during build — the build cache. See below.

Topic 7: The Host Ran Out of Disk

This causes failures that look like everything else — builds failing, containers refusing to start, databases going read-only — so check it early. It is cheap.

df -h /var/lib/docker
docker system df -v          # images, containers, volumes, build cache — with RECLAIMABLE
sudo du -sh /var/lib/docker/*

The usual culprits, in order:

  1. Build cache. On a CI host this grows without bound. docker builder prune is routine maintenance.
  2. Log files. The default json-file driver has no size limit. One chatty container can produce hundreds of gigabytes. Fix it globally in /etc/docker/daemon.json:
{
  "log-driver": "json-file",
  "log-opts": { "max-size": "10m", "max-file": "3" }
}
  1. Exited containers, each still holding its writable layer. docker ps -a shows them; docker container prune clears them.
  2. Anonymous volumes accumulating one per container start (lesson 8).
docker container prune
docker image prune -a
docker builder prune
docker system prune            # the first three at once — safe, does not touch volumes

Handle volumes separately and deliberately. docker system prune --volumes removes volumes not attached to a running container, which includes data you meant to keep. Run docker volume ls and decide, one at a time.

Try it yourself: Fill a container’s log deliberately (docker run -d --name noisy alpine sh -c 'while :; do echo spam; done'), let it run for a minute, then find the log file under /var/lib/docker/containers/<id>/ and check its size. Then restart it with --log-opt max-size=1m and watch the file stop growing. That one default has caused more outages than any exotic failure mode.

Common mistake: Debugging inside the container first. docker exec changes the container, takes the longest, and is often impossible anyway on a distroless image or one that has already exited. docker ps -a, docker logs, docker inspect — in that order, every time. The answer is in one of the three far more often than not.