Containers & Linux Internals
Containers & Linux Internals
Learn containers & linux internals concepts for DevOps interviews
14 — Containers & Linux Internals
The capstone that connects Linux to Docker and Kubernetes. The single most impressive thing you can say in a container interview is “a container is just a Linux process with namespaces and cgroups.” This file makes that true understanding, not a slogan.
1. The big idea: containers are just Linux processes
A container is not a VM. There’s no guest kernel, no hypervisor. A container is a normal process on the host that the kernel has isolated (namespaces) and limited (cgroups), usually with a private root filesystem (image layers).
Virtual Machine Container
┌─────────────┐ ┌─────────────┐
│ App │ │ App │ ← just a host process
│ Guest OS │ full kernel │ (no guest OS / no guest kernel)
│ (kernel) │ └──────┬──────┘
└──────┬──────┘ │ namespaces (isolation) + cgroups (limits)
Hypervisor Host Linux kernel (SHARED)
│ │
Host hardware Host hardware
🎯 Interview signal — say this verbatim: “A container is a host process isolated by namespaces and constrained by cgroups, using an image as its root filesystem. It shares the host kernel — that’s why it’s lighter than a VM and why there’s no guest OS to boot.”
⚠️ Consequence of the shared kernel: containers on a host all use the host’s kernel — you can’t run a Windows container on a Linux kernel, and a kernel exploit can cross the container boundary (weaker isolation than a VM). That trade-off — density/speed vs isolation — is a common follow-up.
2. Namespaces — isolate what a process sees
A namespace virtualizes a kernel resource so processes inside see their own private view. Docker/containerd create a set of namespaces per container.
| Namespace | Isolates | Effect inside the container |
|---|---|---|
| PID | Process IDs | Container’s main process sees itself as PID 1; can’t see host processes |
| NET | Network stack | Own interfaces, IPs, routing table, ports |
| MNT | Mount points | Own filesystem view (its image), own /proc |
| UTS | Hostname | Own hostname |
| IPC | Shared memory / semaphores | Isolated IPC |
| USER | UID/GID mapping | Can be “root” inside but map to an unprivileged UID on the host |
| cgroup | cgroup root view | Sees its own cgroup hierarchy |
lsns # list namespaces on the host
ls -l /proc/<pid>/ns/ # a process's namespaces (symlinks)
unshare --pid --fork --mount-proc bash # manually create a PID namespace (see PID 1!)
🎯 Interview signal: “PID namespace is why the app inside a container is PID 1 and can’t see the host’s processes; NET namespace gives it its own IP/ports; USER namespace lets it be root inside while being unprivileged on the host — a key security control.” Naming specific namespaces and their effect is the depth interviewers want.
⚠️ Because the container’s main process is PID 1, it inherits PID 1’s duties: it must
reap zombies and handle SIGTERM for graceful shutdown. Many app processes don’t — hence
tini/--init as a tiny PID-1 init in containers (ties back to file 03).
3. cgroups — limit what a process can use
Control groups (cgroups) limit and account for a process group’s resources: CPU,
memory, I/O, PIDs. This is how docker run -m 512m --cpus 1.5 and Kubernetes requests/limits
are enforced.
# cgroups v2 (modern) live under /sys/fs/cgroup
cat /sys/fs/cgroup/.../memory.max # memory limit
cat /sys/fs/cgroup/.../cpu.max # CPU quota
systemd-cgls # cgroup tree (systemd manages cgroups too!)
systemd-cgtop # top-like view per cgroup
- Memory limit hit → the cgroup OOM-kills the offending process (container shows
OOMKilled, exit 137 — file 10). It does not kill the whole host. - CPU limit → the process is throttled (not killed) to its quota.
🎯 Interview signal: “cgroups enforce the limits; namespaces enforce the isolation. A container
memory limit is a cgroup memory.max; exceeding it triggers a cgroup OOM kill of that
container (exit 137), not the host.” Splitting the two concepts cleanly is the mark of real
understanding.
⚠️ Note: systemd itself uses cgroups to manage services (file 04) — so cgroups aren’t container-only; they’re a general kernel resource-control mechanism.
4. The container root filesystem: images & layers
A container’s filesystem comes from an image built of read-only layers stacked with a union/overlay filesystem (OverlayFS), plus a thin writable layer on top per container.
Container writable layer (per container, ephemeral) ← changes here vanish when removed
──────────────────────────────────────────────
Image layer 3 (your app) ┐
Image layer 2 (dependencies) │ read-only, SHARED across containers (dedup)
Image layer 1 (base OS userland) ┘
- Layers are cached and shared — many containers from the same image share the read-only layers, saving disk and speeding pulls.
- The writable layer is ephemeral — data written there is lost when the container is removed. Persistent data needs a volume (bind mount or named volume) = a host path mounted in via the MNT namespace.
⚠️ “Base OS userland,” not a kernel: an ubuntu image ships Ubuntu’s userland/libraries but
uses the host kernel. That’s why image “OS” ≠ a real OS install.
5. Putting it together: what docker run actually does
docker run -m 512m --cpus 1 nginx
1. Pull the image layers (if not cached); assemble a root fs via OverlayFS.
2. Create NAMESPACES (PID/NET/MNT/UTS/IPC/USER) → isolated view.
3. Create a CGROUP with memory.max=512M and cpu quota=1 core → resource limits.
4. Set up networking (veth pair into a bridge, or host/none) in the NET namespace.
5. pivot_root/chroot into the image's filesystem (MNT namespace).
6. execve the entrypoint as PID 1 inside the PID namespace.
→ A normal Linux process, isolated and limited. That's the whole "magic."
🎯 Interview signal: being able to narrate this — namespaces for isolation, cgroups for limits, overlay fs for the root, veth for networking, exec as PID 1 — demonstrates you understand containers as Linux primitives, which is exactly what separates strong DevOps candidates.
6. Why this matters for Kubernetes/EKS
- Requests/limits = cgroup settings; a pod over its memory limit is cgroup-OOM-killed
(
OOMKilled, 137). - Pod networking = network namespaces + virtual interfaces (CNI wires them up).
- Security contexts (
runAsNonRoot, drop capabilities, seccomp) = Linux user namespaces, capabilities, and seccomp/AppArmor/SELinux applied to the container process. kubectl exec= entering the container’s namespaces to run a process alongside it.- PID 1 / graceful shutdown = the SIGTERM handling and zombie reaping from file 03, which is why Kubernetes sends SIGTERM then SIGKILL after the grace period.
🎯 This is the bridge to the EKS material: “Kubernetes is an orchestrator on top of these Linux primitives — it schedules processes that the kernel isolates with namespaces and limits with cgroups.”
Interview Questions
Q1. What is a container, really?
A normal Linux process on the host that the kernel isolates with namespaces (so it sees its own PIDs, network, mounts, hostname, users) and constrains with cgroups (CPU, memory, I/O limits), using an image’s layered filesystem as its root. It shares the host kernel — there’s no guest OS or hypervisor — which is why it’s lighter and faster than a VM but has weaker isolation.
Q2. How is a container different from a virtual machine?
A VM runs a full guest OS with its own kernel on a hypervisor, giving strong isolation at the cost of overhead and boot time. A container is just a host process isolated by namespaces and limited by cgroups, sharing the host kernel — so it’s lightweight, starts instantly, and packs densely, but the shared kernel means weaker isolation (a kernel exploit can cross the boundary) and you can’t run a different-kernel OS.
Q3. What do namespaces do? Name a few.
Namespaces isolate what a process can see by virtualizing kernel resources. PID (own process tree; the app is PID 1 and can’t see host processes), NET (own interfaces/IPs/ports), MNT (own filesystem view), UTS (own hostname), IPC (own shared memory), and USER (map container-root to an unprivileged host UID). Each container gets its own set.
Q4. What do cgroups do, and what happens when a memory limit is exceeded?
cgroups limit and account for a process group’s resources — CPU, memory, I/O, PID count. When a container exceeds its memory limit, the kernel OOM-kills the offending process within that cgroup (the container shows OOMKilled, exit 137), without taking down the host. A CPU limit throttles the process rather than killing it. Namespaces isolate; cgroups limit.
Q5. Why is the container’s process PID 1, and why does that matter?
The PID namespace makes the container’s entrypoint PID 1 in its own process tree. That matters because PID 1 has special duties — reaping zombie children and handling signals. Many app processes don’t reap zombies or handle SIGTERM for graceful shutdown, so containers often run a tiny init (
tini/--init) as PID 1 to do that. It’s the same PID-1 responsibility as on a full system.
Q6. How does a container’s filesystem work?
From an image made of read-only layers stacked with a union/overlay filesystem, plus a thin writable layer per container. Read-only layers are cached and shared across containers for efficiency. The writable layer is ephemeral — data written there is lost when the container is removed — so persistent data uses volumes (host paths mounted in via the mount namespace). The base image ships userland, not a kernel; it uses the host’s.
Q7. A container was OOMKilled with exit code 137. Explain the whole chain.
The container’s cgroup had a memory limit; the process tried to use more than
memory.max, so the kernel’s OOM killer terminated it within that cgroup. 137 = 128 + 9, i.e., killed by SIGKILL. It’s confined to the container — the host and other containers keep running. The fix is raising the limit or fixing the memory leak/right-sizing, and I’d confirm via the container runtime’s status and the host’s dmesg/journal.
Q8. (Senior) Connect containers to how Kubernetes works.
Kubernetes is an orchestrator over these Linux primitives. Pod resource requests/limits map to cgroup settings (over-limit = cgroup OOM kill, 137); pod networking is network namespaces wired by a CNI with virtual interfaces; security contexts use user namespaces, Linux capabilities, and seccomp/AppArmor/SELinux on the container process;
kubectl execenters the container’s namespaces; and graceful shutdown relies on PID 1 handling SIGTERM, which is why the kubelet sends SIGTERM and then SIGKILL after the termination grace period. Understanding containers as namespaces + cgroups + overlay fs is exactly what makes Kubernetes behavior predictable.
That’s the full concept curriculum. Now prove it:
- Master Interview Questions — cross-cutting scenarios.
- Projects — hands-on builds that combine everything.