Performance Monitoring
Performance Monitoring
Learn performance monitoring concepts for DevOps interviews
10 — Performance & Resource Monitoring
“The server is slow” is the most common on-call page. Interviewers want a methodical answer: is it CPU, memory, disk I/O, or network? This file gives you the mental model (USE method), the tools, and the two facts people always get wrong: load average and the OOM killer.
1. The four resources & the USE method
Any performance problem is one (or a mix) of: CPU, Memory, Disk I/O, Network.
The USE method: for each resource check Utilization, Saturation, Errors.
- Utilization: how busy (e.g., CPU %). Saturation: how much queued/waiting (run queue, swap). Errors: dropped packets, disk errors.
🎯 Interview signal: framing “server is slow” as “I check CPU, memory, disk, and network for
utilization and saturation, then drill into the bottleneck” beats randomly running top.
2. Load average — the most misunderstood metric
uptime
# load average: 4.20, 3.80, 2.10 ← 1-min, 5-min, 15-min
cat /proc/loadavg
⚠️ Load average is NOT CPU percentage. On Linux it’s the average number of processes in the run queue — those running (R) plus those in uninterruptible sleep (D) (waiting on I/O). So:
- Compare load to core count: load
4.0on a 4-core box ≈ fully busy; on an 8-core box ≈ half-loaded.nprocgives the core count. - High load + low CPU% → the queue is full of I/O-blocked (D-state) tasks, not CPU work. That points to disk/NFS, not CPU.
- Rising 1-min vs 15-min average = the problem is getting worse.
🎯 Interview signal: “Linux load average includes uninterruptible I/O-wait tasks, so high load
with idle CPUs means an I/O bottleneck — I’d look at iostat/iowait, not add CPU.” This one
answer separates seniors from juniors.
3. CPU
top # live: %us (user), %sy (system/kernel), %id (idle), %wa (I/O WAIT), %st (steal)
htop # nicer top
mpstat -P ALL 1 # per-core utilization
pidstat 1 # per-process CPU over time
Reading top’s CPU line:
- %us high → application/user code busy (real CPU work).
- %sy high → lots of kernel/syscalls (context switches, I/O).
- %wa high → CPU waiting on disk I/O (not a CPU problem — a storage one).
- %st (steal) high → the hypervisor is giving your vCPU to other tenants (noisy neighbor / oversubscribed VM). ⚠️ On cloud VMs, high steal = you need a bigger/less-contended instance; it’s not your app.
4. Memory — and why free confuses people
free -h
# total used free shared buff/cache available
# Mem: 16Gi 4Gi 1Gi 0.2Gi 11Gi 11Gi
⚠️ “free” being low is normal and good. Linux uses spare RAM for buff/cache (disk
cache) to speed things up, and reclaims it instantly when apps need memory. The number that
matters is available — how much apps can actually get. “My server only has 1G free!” is
usually fine if available is high.
vmstat 1 # si/so columns = swap in/out — nonzero = SWAPPING (bad)
cat /proc/meminfo
ps aux --sort=-%mem | head # top memory consumers
- Swapping (
si/sononzero, or high%wa) = out of RAM, pages going to disk → performance collapse (“thrashing”).
5. The OOM killer (guaranteed question)
When memory is truly exhausted (RAM + swap), the kernel’s Out-Of-Memory killer picks a
process to kill to save the system. It scores processes (oom_score, roughly by memory use
and oom_score_adj) and kills the highest.
dmesg | grep -i "killed process" # OOM kill evidence
journalctl -k | grep -i oom
cat /proc/<pid>/oom_score
⚠️ “My process died with no error / exit 137” → often the OOM killer (or a SIGKILL; 137 =
128 + 9). In containers/Kubernetes, exceeding the memory limit triggers a cgroup OOM kill
of that container (OOMKilled). The fix is right-sizing memory or fixing a leak, not just
restarting.
🎯 Interview signal: “Exit code 137 = 128 + SIGKILL(9), classically the OOM killer or a cgroup
memory-limit kill. I’d check dmesg/journalctl -k for the OOM message and the process’s
oom_score.” Naming 137 = OOM is a strong signal.
6. Disk I/O
iostat -xz 1 # per-device: %util, await (latency), r/s w/s, aqu-sz (queue)
iotop # per-process I/O (needs root)
sudo dstat # combined cpu/disk/net/memory live view
Key columns in iostat -x:
- %util near 100% → the disk is saturated.
- await (ms) high → I/O latency is bad (slow disk, overloaded volume).
- High disk I/O also shows as %wa in top and D-state processes.
⚠️ On cloud, EBS-type volumes have IOPS/throughput limits; hitting them looks like a slow
disk even though the “disk” is fine — you’ve exhausted provisioned IOPS. Match iostat to the
volume’s limits.
7. Network (perf angle)
ss -s # socket summary
ip -s link # interface errors/drops
sar -n DEV 1 # per-interface throughput over time
iftop / nload # live bandwidth
High drops/errors on an interface, or saturated bandwidth, point to a network bottleneck (file 06 covers connectivity; this is throughput/saturation).
8. The “server is slow” playbook (say this)
1. uptime / load average → compare to nproc. Is load high? getting worse?
2. top → is it %us (CPU-bound app), %wa (I/O-bound), %sy (kernel), %st (noisy neighbor)?
3. free -h / vmstat → is 'available' low? any swapping (si/so)?
4. iostat -xz 1 → is a disk at 100% %util with high await?
5. Identify the culprit process: pidstat / iotop / ps --sort.
6. Correlate with logs (journalctl, app logs) and recent changes/deploys.
7. dmesg → OOM kills? hardware/driver errors?
🎯 Interview signal: a structured walk (resource by resource, top-down) with the reasoning at each step (“high %wa means I go to iostat, not add CPU”) is exactly what interviewers score.
Interview Questions
Q1. What does load average actually measure?
The average number of processes in the run queue — those running or runnable plus, on Linux, those in uninterruptible sleep (D state, usually waiting on disk/NFS I/O). It’s not a CPU percentage. You interpret it relative to core count: load equal to the number of cores is roughly fully utilized. Rising 1-minute vs 15-minute values mean the situation is worsening.
Q2. Load is very high but CPU usage is low. What’s happening?
The run queue is full of I/O-blocked (D-state) tasks, not CPU work — an I/O bottleneck. Because Linux includes uninterruptible-sleep tasks in load, a hung/slow disk or NFS mount inflates load while CPUs sit idle. I’d look at
iostat -xfor %util/await, top’s %wa, and which processes are in D state — not add CPU.
Q3. free -h shows almost no free memory. Is that a problem?
Usually not. Linux uses otherwise-idle RAM for buffer/page cache to speed up I/O and reclaims it instantly when applications need memory. The meaningful number is available, not free. A problem exists only if
availableis low and the system is swapping (nonzero si/so invmstat), which indicates real memory pressure.
Q4. A process disappeared with exit code 137 and no application error. Why?
137 = 128 + 9, i.e., killed by SIGKILL — classically the kernel OOM killer reclaiming memory when the system (or a cgroup/container memory limit) ran out. I’d confirm with
dmesgorjournalctl -kfor the “Killed process” / “Out of memory” message and check the process’soom_score. Fix is right-sizing memory or fixing a leak, not just restarting.
Q5. Walk me through diagnosing “the server is slow.”
Start with
uptimeto see load vs core count and trend. Thentopto classify: %us = CPU-bound app, %wa = I/O wait, %sy = kernel-heavy, %st = hypervisor steal on a VM. Check memory withfree/vmstat(isavailablelow, is it swapping?). If I/O-bound,iostat -xz 1for %util/await, andiotop/pidstatto find the process. Correlate with logs, recent deploys, anddmesgfor OOM/hardware errors. It’s resource-by-resource, top-down, with reasoning at each step.
Q6. What does high %st (steal) in top mean?
On a virtual machine, steal time is CPU the hypervisor gave to other tenants while your vCPU was ready to run — a sign of an oversubscribed host or a noisy neighbor. Your app isn’t the problem; the fix is a larger/less-contended instance or moving off the busy host.
Q7. How do you tell a disk-I/O bottleneck from a CPU bottleneck?
CPU-bound shows high %us with low %wa and the app on-CPU (pidstat). I/O-bound shows high %wa, D-state processes, and in
iostat -xa device near 100% %util with high await. Load average can be high in both, so I distinguish them by whether time is spent on-CPU vs waiting on I/O.
Q8. (Senior) A cloud VM’s disk seems slow but the disk hardware is healthy. What do you suspect?
Provisioned IOPS/throughput exhaustion. Cloud block volumes (EBS-style) have per-volume IOPS and throughput limits (and burst credits on some tiers). If the workload exceeds them,
iostatshows high await/%util even though nothing is “broken.” I’d compare observed IOPS/ throughput to the volume’s provisioned limits and either raise them, switch volume type, use a larger volume, or spread I/O — and check for burst-credit depletion on smaller volumes.