Linux Architecture & Boot Process

0%
20 minarchitecturebootkernel

Linux Architecture & Boot Process

Learn linux architecture & boot process concepts for DevOps interviews

00 — Linux Architecture & Boot Process

You know how to use Linux. This file gives you the mental model interviewers expect: what the kernel is, how userspace talks to it, and exactly what happens from power-on to login prompt. Every later topic (processes, permissions, systemd) hangs off this.


1. Kernel space vs user space (the root mental model)

┌──────────────────────────────────────────────────────────────┐
│                        USER SPACE                              │
│   your shell, nginx, python, systemd services, ssh...          │
│   - can't touch hardware directly                              │
│   - asks the kernel for everything via SYSTEM CALLS            │
└───────────────────────────────┬───────────────────────────────┘
                                │  syscalls: open, read, write,
                                │  fork, execve, socket, mmap...
┌───────────────────────────────▼───────────────────────────────┐
│                       KERNEL SPACE                             │
│   process scheduler, memory manager, filesystems, network      │
│   stack, device drivers                                        │
└───────────────────────────────┬───────────────────────────────┘
                                ▼
                            HARDWARE (CPU, RAM, disk, NIC)
  • Kernel: privileged core that manages CPU scheduling, memory, filesystems, networking, and devices. Runs in a protected CPU mode.
  • User space: everything else — your programs. They can’t access hardware directly; they request services through system calls.
  • System call: the controlled doorway into the kernel (e.g., read(), write(), fork(), execve(), socket()). Libraries like glibc wrap them.

🎯 Interview signal: “User programs never touch hardware directly — they make system calls, and the CPU switches to kernel mode to service them. That boundary is what keeps the system safe and multi-tenant.” Trace a command with strace ls to see the syscalls.

⚠️ Don’t confuse the kernel with the OS/distribution. The kernel is Linux itself; Ubuntu/RHEL are distributions = kernel + GNU tools + package manager + defaults.


2. What “Linux” actually includes

Layer Examples
Kernel Process scheduling, memory, VFS, network stack, drivers
Core userland (GNU) bash, coreutils (ls, cp), glibc
Init system systemd (PID 1) on most modern distros
Service/daemons sshd, nginx, cron, chronyd
Package manager apt/dpkg (Debian family), dnf/rpm (RHEL family)

3. The boot sequence (power-on → login)

Be able to recite this end to end — it’s a very common question.

1. FIRMWARE (BIOS/UEFI)
   - Power-on self test, finds a bootable device.
2. BOOTLOADER (GRUB2)
   - Loads the Linux KERNEL + INITRAMFS into memory, passes kernel cmdline args.
3. KERNEL
   - Initializes CPU, memory, drivers; mounts a temporary root from INITRAMFS.
4. INITRAMFS (initial RAM filesystem)
   - Tiny temporary root with just enough drivers/modules to find and mount the REAL root fs.
5. Kernel mounts the real root filesystem, then starts PID 1.
6. PID 1 = systemd (init)
   - Brings the system to a "target" (set of services), starts everything in dependency order.
7. Reaches default target (e.g., graphical.target or multi-user.target)
   - getty/display manager shows a login prompt. sshd is up. System is ready.

Key artifacts:

  • GRUB config: /boot/grub/grub.cfg (generated; edit /etc/default/grub + update-grub).
  • Kernel + initramfs: live in /boot.
  • Kernel command line: cat /proc/cmdline shows what the bootloader passed (root device, quiet, ro, etc.).

🎯 Interview signal: naming initramfs and why it exists (a chicken-and-egg fix: the kernel needs drivers/modules to mount the real root, but those may live on that root, so a tiny in-RAM root bootstraps it) is a strong senior signal.

⚠️ “Server won’t boot” classics: broken /etc/fstab (drops you to emergency shell), a bad GRUB entry, or a kernel/initramfs mismatch after a botched upgrade. You recover from the GRUB menu (pick an older kernel) or a rescue/live environment.


4. systemd as PID 1 (the init system)

The kernel starts exactly one process directly: PID 1, which on modern Linux is systemd. It is the ancestor of every other process and the service manager.

What PID 1 / systemd does:

  • Starts services in dependency order (parallelized, not a serial script).
  • Reaps orphaned processes (adopts children whose parent died — see file 03).
  • Manages the system’s state via targets and keeps services running (restart policy).
  • Owns the journal (logging, file 11) and timers (scheduling, file 13).

⚠️ If PID 1 dies, the kernel panics — the system halts. That’s why PID 1 is special.

The old vs new

  • SysVinit (legacy): serial shell scripts in /etc/init.d, numbered “runlevels.”
  • systemd (modern): parallel, dependency-aware, declarative unit files. Runlevels are replaced by targets.

5. Targets (the modern “runlevels”)

A target is a named group of units representing a system state.

Target Old runlevel Meaning
poweroff.target 0 Halt
rescue.target 1 Single-user/maintenance
multi-user.target 3 Full system, no GUI (typical server)
graphical.target 5 Multi-user + GUI
reboot.target 6 Reboot
systemctl get-default                 # what target boots by default
systemctl set-default multi-user.target   # servers: no GUI
systemctl isolate rescue.target       # switch to maintenance now
systemctl list-units --type=target    # active targets

🎯 Interview signal: “Servers boot to multi-user.target (no GUI) to save resources; graphical.target pulls in multi-user.target plus a display manager.” Understanding that targets have dependencies (one pulls in another) shows depth.


6. Inspecting the running system

uname -r                     # kernel version
cat /etc/os-release          # distro + version
systemctl status             # overall system state, PID 1
ps -p 1 -o comm=             # confirm PID 1 is systemd
systemd-analyze              # how long boot took
systemd-analyze blame        # which units were slowest to start
systemd-analyze critical-chain  # the dependency chain that gated boot time
journalctl -b                # logs for the current boot

⚠️ systemd-analyze blame is the go-to for “boot is slow” questions — it ranks units by startup time so you can find the culprit (often a service waiting on network/DNS).


Interview Questions

Q1. Explain the difference between kernel space and user space.

The kernel runs in a privileged CPU mode with full hardware access; it manages scheduling, memory, filesystems, networking, and drivers. User space is where normal programs run — they can’t touch hardware directly and must request services through system calls (open, read, fork, execve). The CPU switches to kernel mode to service a syscall, then returns. That boundary provides isolation and security.

Q2. Walk me through the Linux boot process.

Firmware (BIOS/UEFI) POSTs and finds a boot device; the bootloader (GRUB2) loads the kernel and initramfs into memory; the kernel initializes hardware and mounts the initramfs as a temporary root; initramfs loads the drivers needed to mount the real root filesystem; the kernel then starts PID 1 (systemd); systemd brings the system to its default target, starting services in dependency order until you get a login prompt with sshd up.

Q3. What is initramfs and why does it exist?

It’s a small temporary root filesystem loaded into RAM by the bootloader. It solves a chicken-and-egg problem: the kernel may need drivers/modules (for the disk controller, LVM, RAID, encryption) to mount the real root filesystem, but those modules live on that root. initramfs contains just enough to find and mount the real root, then hands off.

Q4. What is PID 1 and what happens if it dies?

PID 1 is the first userspace process the kernel starts — systemd on modern distros. It’s the ancestor of all processes, the service manager, and the reaper of orphaned processes. If PID 1 exits, the kernel panics and the system halts, which is why it’s treated specially.

Q5. What replaced runlevels in systemd, and which does a server use?

Targets. A target is a named group of units representing a system state. A typical server boots to multi-user.target (full multi-user, no GUI) instead of graphical.target to save resources. Targets have dependencies — graphical pulls in multi-user plus a display manager.

Q6. A server boots slowly. How do you find out why?

systemd-analyze for total boot time, systemd-analyze blame to rank units by startup duration, and systemd-analyze critical-chain to see the dependency chain that actually gated boot. Slow units are often services waiting on the network or DNS; I’d check that unit’s journal (journalctl -u <unit> -b).

Q7. How is the Linux kernel different from a Linux distribution?

The kernel is Linux itself — the core that manages hardware and processes. A distribution is the kernel plus userland (GNU coreutils, bash, glibc), an init system (systemd), a package manager (apt/dnf), and defaults. Ubuntu and RHEL ship different userland/tooling around the same kind of kernel.

Q8. (Senior) A box drops to an emergency shell on boot after a change. What’s your first suspect?

A broken /etc/fstab — an unmountable or missing filesystem entry makes systemd fail the local-fs.target and drop to emergency mode. I’d read the journal/console error, boot with the root fs read-write from the emergency shell (or a rescue environment), fix or comment the bad fstab line, and remount. Other suspects: a bad GRUB entry or a kernel/initramfs mismatch from an interrupted upgrade (boot an older kernel from GRUB).


Next: 01 — Filesystem & FHS — how the filesystem is laid out and what inodes really are.