Namespaces and cgroups: What Isolation Really Is

The two kernel features that make a container a container — which namespaces exist, what each one hides, and why cgroups are the only thing standing between one container and the whole host.

beginner 20 min lesson hands-on task included

There is no such thing as a container in the Linux kernel. There is no container syscall and no container object. What exists is a process with namespaces applied to restrict its view and cgroups applied to restrict its consumption. “Container” is the name we give that combination.

Knowing this is not academic. It is why --pid=host works, why a container can see the host’s /proc when you mount it, and why “container escape” is a category of vulnerability rather than a contradiction in terms.


Topic 1: Process Isolation, From First Principles

A process is a program in execution. Every process gets memory, CPU time from the scheduler, and a process ID. That is the ordinary Linux model.

Container technology on Linux predates Docker. LXC (Linux Containers) did OS-level virtualisation — multiple isolated Linux systems on one kernel — years before Docker existed. Docker’s original contribution was not the isolation primitives. It was making them easy: a packaging format, a build file, a registry, and one command.

What a container gets, and why it feels like a machine:

The container hasProvided by
Its own process treePID namespace
Its own network interface and IPNetwork namespace
Its own filesystem view, rooted at /Mount namespace
Its own hostnameUTS namespace
Its own users and groupsUser namespace
Its own shared memory and semaphoresIPC namespace
A share of CPU, RAM and disk I/Ocgroups

From the host’s point of view, all of that is still just one process among many. From inside, it looks like a complete operating system. Both views are correct.


Topic 2: The Namespaces, One by One

PID namespace gives the container a virtual process tree. Its main process is PID 1 inside; on the host it has an ordinary PID like 24817. The nesting is one-way: the host sees every process inside every container, but a container cannot see the host’s processes or its siblings.

docker exec iso ps aux              # sleep is PID 1
ps -ef | grep 'sleep 1d'            # on the host: same process, a different PID

Being PID 1 has a real consequence. PID 1 in a namespace does not get default signal handlers, and it inherits responsibility for reaping orphaned children. An application that ignores SIGTERM because it never installed a handler will not stop on docker stop — it will sit there until the ten-second grace period expires and SIGKILL arrives. Init systems designed for containers (tini, dumb-init, or docker run --init) exist entirely to fix this.

Network namespace gives each container its own network stack: interfaces, routing table, iptables rules, DNS configuration. This is why every container gets its own IP address and why two containers can both listen on port 80 without colliding. Under the hood, a virtual ethernet pair is created — one end inside the container as eth0, the other on the host (you will see it in ifconfig as something like vethf5e485a), plugged into a bridge.

Mount namespace gives each container its own view of the mount table. This is what makes the image’s filesystem appear as /. The host’s /etc and the container’s /etc are unrelated files.

UTS namespace isolates hostname and domain name — which is why hostname inside a container returns the container ID by default.

IPC namespace isolates System V IPC and POSIX message queues, so shared-memory segments in one container are invisible to another.

User namespace allows a whole separate set of users and groups, and crucially allows UID mapping: root (UID 0) inside the container can map to an unprivileged UID on the host. This is the foundation of rootless Docker. It is worth noting that user namespace remapping is not enabled by default — by default, root in the container is root on the host, minus some capabilities. Lesson 12 takes this apart properly.

You can inspect all of this directly, because the kernel exposes it:

ls -l /proc/self/ns                     # your shell's namespaces, as inode numbers
sudo ls -l /proc/<container-pid>/ns     # the container's — different inodes = different namespaces

Two processes in the same namespace share the same inode number for it. That single comparison is the most convincing demonstration of what a container is that you will ever run.


Topic 3: Sharing Namespaces on Purpose

Namespaces are opt-out as well as opt-in, and the opt-outs are genuinely useful for debugging:

docker run --rm -it --pid=host alpine ps aux           # see every host process
docker run --rm -it --net=host nginx                   # use the host's network stack directly
docker run --rm -it --net=container:api nicolaka/netshoot  # join another container's network
docker run --rm -it --uts=host alpine hostname         # share the host's hostname

--net=container:<name> is the one to remember. It puts a debugging container inside the target’s network namespace, so tcpdump, ss and curl see exactly what the application sees — without installing any of those tools into your production image. (If Kubernetes pods feel familiar here: a pod is a group of containers sharing exactly this set of namespaces.)

Each of these weakens isolation in a specific, bounded way. --pid=host in production is a red flag on a code review. --net=host is common and often justified for performance-sensitive network tooling, and it means the container can bind directly to host ports with no publishing and no isolation.


Topic 4: cgroups — Limits, Not Visibility

Namespaces control what a process can see. Control groups (cgroups) control what it can use. They are a separate kernel feature and they answer a separate question.

Without cgroups, isolation is a security story with no capacity story: a container would still be able to consume every core and all the RAM on the box.

cgroups impose limits on:

  • CPU — as a quota (a share of time per period) or as a CPU set (specific cores)
  • Memory — a hard ceiling, enforced by the OOM killer
  • Block I/O — read/write bandwidth and IOPS per device
  • PIDs — maximum number of processes, which is the fork-bomb defence
docker run -d --cpus 1 --memory 128m --pids-limit 100 --name gov gol:1.0
docker inspect gov --format '{{.HostConfig.Memory}} {{.HostConfig.NanoCpus}}'
docker stats gov                        # the runtime view of both

Modern distributions run cgroup v2, which unifies the controller hierarchy and changes some of the accounting. docker info tells you which version you are on, and it is worth knowing before you argue with a memory graph.

The gotcha that produces mysterious OOMs:

Many runtimes read /proc/meminfo to size themselves — and /proc/meminfo is not namespaced. A JVM or a Node process inside a 512 MB container may see the host’s 64 GB and set its heap accordingly, then get OOM-killed the moment it tries to use it. Modern JVMs are container-aware (-XX:+UseContainerSupport is on by default in current versions); older runtimes and many language libraries are not. If a process is being killed at a memory level far below what it “thinks” it has, this is almost always why.


Topic 5: What This Means for Security

Now the honest summary of the trade-off from lesson 1, with the mechanism visible.

A VM boundary is enforced by the hypervisor and the CPU’s virtualisation extensions. A container boundary is enforced by kernel data structures, in a kernel that both the container and the host share. Every syscall a container makes is handled by your kernel.

That has direct consequences:

  • A kernel privilege-escalation bug is a container escape.
  • --privileged disables essentially all of the restrictions — capabilities, device access, seccomp — and should be treated as “run this as root on the host”.
  • Mounting /var/run/docker.sock into a container gives it control of the daemon, and the daemon can start a container that mounts the host’s root filesystem. Same outcome, no exploit required.
  • Root inside a container is root on the host unless user-namespace remapping or rootless mode is in use.

None of this means containers are insecure. It means the isolation is a kernel-enforced boundary, not a hardware one, and the defences that matter — dropping capabilities, non-root users, seccomp, read-only root filesystems — are the subject of lesson 12.

Try it yourself: Run docker run --rm -it --privileged alpine sh and then fdisk -l. You are looking at the host’s disks from inside a container. Now run the same without --privileged and watch it show nothing. That is the entire security model in two commands.

Common mistake: Believing a container limit is a limit the application can see. It is a limit the kernel enforces. Your application will happily plan to use memory it will never be allowed to have. Configure the runtime’s own limits (JVM heap, Node --max-old-space-size, worker counts) to match the cgroup limit rather than assuming the two agree.