Processes & Signals

0%
20 minprocessessignalskernel

Processes & Signals

Learn processes & signals concepts for DevOps interviews

03 — Processes & Signals

Processes and signals are the beating heart of Linux ops. Interviewers test whether you understand what a process really is, how new ones are created (fork/exec), the process states (especially D and Z), and how signals like SIGTERM vs SIGKILL differ.


1. Program vs process

  • A program is a file on disk (an executable).
  • A process is a running instance: it has a PID, a parent (PPID), its own memory (address space), open file descriptors, environment, and a UID/GID it runs as.
  • A thread is a unit of execution inside a process sharing its memory. (Linux implements threads as processes that share resources — “tasks”.)
ps aux                 # all processes (BSD style): USER PID %CPU %MEM ... STAT ... COMMAND
ps -ef                 # all processes (System V style): UID PID PPID ... CMD
ps -eLf                # include threads (LWP column)
pstree -p              # process hierarchy with PIDs
top    /  htop         # live view
pgrep -a nginx         # find PIDs by name

2. How a new process is created: fork() + execve()

The classic model (and a favorite question):

Parent process
   │  fork()          → creates a CHILD that is a near-exact COPY of the parent
   ▼                    (same code, copied memory via copy-on-write, new PID)
Child process
   │  execve("/bin/ls")→ REPLACES the child's memory image with a new program
   ▼
Now the child is running `ls`; parent can wait() for it to finish.
  • fork() duplicates the calling process (copy-on-write memory, so it’s cheap). Returns the child PID to the parent and 0 to the child.
  • execve() replaces the current process image with a new program (same PID).
  • wait()/waitpid() lets the parent collect the child’s exit status (reaping).

🎯 Interview signal: “The shell forks a child, the child execs the command, and the shell waits for it. That fork-then-exec pattern is how every command you run starts.” Show it: strace -f -e trace=fork,execve bash -c 'ls'.


3. Process states (know D and Z)

The STAT column in ps:

State Meaning
R Running or runnable (on the run queue)
S Interruptible sleep (waiting for an event, e.g., network) — normal idle state
D Uninterruptible sleep — usually blocked on I/O; can’t be killed until it returns
T Stopped (e.g., by SIGSTOP/Ctrl-Z) or traced
Z Zombie — finished, waiting for its parent to reap its exit status

Suffixes: s (session leader), + (foreground group), l (multi-threaded), </N (high/low priority).

⚠️ D-state (uninterruptible) tasks can’t be killed, not even by SIGKILL, until the kernel operation they’re blocked in (often disk/NFS I/O) completes. Many D-state processes = an I/O problem (slow disk, hung NFS), and they inflate the load average (file 10). This is a top senior question.


4. Zombies and orphans (must-know pair)

  • Zombie (Z): a child has exited but the parent hasn’t wait()ed to read its exit status, so a tiny entry lingers in the process table. Zombies use no CPU/RAM — just a table slot. The fix is to make the parent reap (or restart/kill the parent); you can’t kill a zombie (it’s already dead).
  • Orphan: a child whose parent died first. It’s re-parented to PID 1 (systemd), which reaps it when it exits. Orphans are normal and harmless.

🎯 Interview signal: “You can’t kill a zombie because it’s already terminated — you kill or fix the parent that failed to reap it. Orphans, by contrast, get adopted by PID 1 and cleaned up automatically.” Naming that distinction cleanly is a strong signal.

⚠️ A pile of zombies indicates a buggy parent not reaping children (common in poorly written app servers) — the real fix is in that parent process.


5. Signals — how you talk to processes

A signal is an asynchronous notification sent to a process. Common ones:

Signal Number Default action Catchable? Typical use
SIGHUP 1 terminate yes Terminal hangup; many daemons reload config on HUP
SIGINT 2 terminate yes Ctrl-C
SIGKILL 9 terminate NO Force kill — cannot be caught/ignored/blocked
SIGTERM 15 terminate yes Polite “please shut down” — the default of kill
SIGSTOP 19 stop NO Pause a process (Ctrl-Z sends SIGTSTP, which is catchable)
SIGCONT 18 continue yes Resume a stopped process
SIGSEGV 11 core dump yes Invalid memory access (crash)
kill <pid>              # sends SIGTERM (15) by default — graceful
kill -TERM <pid>        # explicit graceful
kill -9 <pid>           # SIGKILL — last resort, no cleanup
kill -HUP <pid>         # ask a daemon to reload config
pkill -f "python app"   # by name/pattern
killall nginx           # by exact process name

🎯 Interview signal — the canonical answer: SIGTERM asks the process to shut down cleanly (flush buffers, close files, finish requests); SIGKILL is handled by the kernel and destroys the process immediately with no chance to clean up. Always try SIGTERM first; use SIGKILL only if it won’t exit. SIGKILL and SIGSTOP are the two signals that cannot be caught or ignored.

⚠️ Force-killing (-9) a database or a process mid-write can corrupt data or leave locks — that’s why graceful SIGTERM matters. Kubernetes uses exactly this: it sends SIGTERM, waits terminationGracePeriodSeconds, then SIGKILL.


6. Foreground, background, and jobs

long_task &            # run in background
jobs                   # list shell jobs
fg %1                  # bring job 1 to foreground
bg %1                  # resume job 1 in background
Ctrl-Z                 # suspend the foreground job (SIGTSTP)
nohup long_task &      # keep running after the shell/SSH session closes
disown -h %1           # detach a job from the shell so SIGHUP won't kill it

⚠️ Background jobs started in an SSH session normally die when the session ends (they get SIGHUP). Use nohup, disown, setsid, or better a systemd service / tmux for long-running work. Interviewers love “how do you keep a process running after you log out?”


7. Priorities: nice and renice

The scheduler shares CPU by priority. niceness ranges from −20 (highest priority) to +19 (lowest); default 0. Only root can set negative (higher-priority) values.

nice -n 10 ./batch_job.sh      # start with lower priority (nicer to others)
renice -n 5 -p <pid>           # change a running process's niceness

For I/O priority there’s ionice. 🎯 “A batch job is starving interactive work” → renice it higher (bigger nice value) or use cgroups (file 14) to cap it.


8. Inspecting a process deeply

cat /proc/<pid>/status      # state, memory, uids, threads
ls -l /proc/<pid>/fd        # open file descriptors (and deleted-but-open files)
cat /proc/<pid>/limits      # ulimits in effect
strace -p <pid>             # trace syscalls of a running process (why is it stuck?)
lsof -p <pid>               # open files/sockets

🎯 Interview signal: “If a process is hung, strace -p <pid> shows the syscall it’s stuck in — often a read/write on a slow fd or a futex (lock) — which tells me if it’s I/O, network, or a deadlock.”


Interview Questions

Q1. What’s the difference between a program and a process?

A program is an executable file on disk. A process is a running instance of it with a PID, a parent PID, its own memory/address space, open file descriptors, environment, and the UID it runs as. One program can have many processes running from it.

Q2. How is a new process created in Linux?

With fork() then execve(). fork() duplicates the calling process (copy-on-write memory, new PID); the child then calls execve() to replace its memory image with a new program. The parent typically wait()s to collect the child’s exit status. That fork-then-exec pattern is how the shell runs every command.

Q3. Explain SIGTERM vs SIGKILL. When do you use each?

SIGTERM (15) is a polite request to shut down that the process can catch to clean up — flush buffers, close files, finish in-flight work. SIGKILL (9) is handled by the kernel and destroys the process immediately with no cleanup, and cannot be caught, ignored, or blocked. Always try SIGTERM first; use SIGKILL only if the process won’t exit. Force-killing can corrupt data or leave locks.

Q4. What is a zombie process and how do you get rid of it?

A zombie is a child that has exited but whose parent hasn’t wait()ed to read its exit status, so a slot lingers in the process table. It uses no CPU/RAM. You can’t kill it — it’s already dead. The fix is to make the parent reap it, or restart/kill the parent. Many zombies indicate a buggy parent not reaping its children.

Q5. What’s an orphan process?

A process whose parent terminated before it did. The kernel re-parents it to PID 1 (systemd), which reaps it when it exits. Orphans are normal and harmless — unlike zombies, they get cleaned up automatically.

Q6. A process is in D state and kill -9 won’t remove it. Why?

D is uninterruptible sleep — the process is blocked inside a kernel operation, almost always I/O (disk or a hung NFS mount). It can’t handle any signal, including SIGKILL, until that operation returns. So the fix isn’t more signals — it’s resolving the underlying I/O problem (recover the storage/NFS server). Many D-state tasks also inflate the load average.

Q7. How do you keep a process running after you close your SSH session?

The session close sends SIGHUP to its jobs. Use nohup cmd &, disown an existing job, setsid, or run it in tmux/screen. For anything real/long-lived, the right answer is a systemd service so it’s managed, logged, and restarted.

Q8. What is niceness and when would you change it?

Niceness (−20 to +19, default 0) biases the scheduler’s CPU share; higher nice = lower priority. I’d raise a batch job’s niceness (nice/renice) so it yields to interactive or latency-sensitive work, or lower a critical process’s value (root only). For hard limits I’d use cgroups.

Q9. (Senior) A production app spawns thousands of zombies over a day. Walk me through it.

The parent process isn’t reaping its children — it forks workers but never calls wait()/waitpid() (or doesn’t handle SIGCHLD). Zombies themselves are cheap, but they exhaust PIDs/process-table slots over time, eventually blocking new process creation. Short term: restart the parent (its children get re-parented to PID 1 and reaped). Real fix: patch the parent to reap children or handle SIGCHLD; in containers, run a proper init (e.g., tini) as PID 1 to reap.


Next: 04 — systemd & Service Management — how services are actually run and kept alive.