Performance Monitoring

0%
20 minperformancemonitoringprofiling

Performance Monitoring

Learn performance monitoring concepts for DevOps interviews

10 — Performance & Resource Monitoring

“The server is slow” is the most common on-call page. Interviewers want a methodical answer: is it CPU, memory, disk I/O, or network? This file gives you the mental model (USE method), the tools, and the two facts people always get wrong: load average and the OOM killer.


1. The four resources & the USE method

Any performance problem is one (or a mix) of: CPU, Memory, Disk I/O, Network.

The USE method: for each resource check Utilization, Saturation, Errors.

  • Utilization: how busy (e.g., CPU %). Saturation: how much queued/waiting (run queue, swap). Errors: dropped packets, disk errors.

🎯 Interview signal: framing “server is slow” as “I check CPU, memory, disk, and network for utilization and saturation, then drill into the bottleneck” beats randomly running top.


2. Load average — the most misunderstood metric

uptime
# load average: 4.20, 3.80, 2.10   ← 1-min, 5-min, 15-min
cat /proc/loadavg

⚠️ Load average is NOT CPU percentage. On Linux it’s the average number of processes in the run queue — those running (R) plus those in uninterruptible sleep (D) (waiting on I/O). So:

  • Compare load to core count: load 4.0 on a 4-core box ≈ fully busy; on an 8-core box ≈ half-loaded. nproc gives the core count.
  • High load + low CPU% → the queue is full of I/O-blocked (D-state) tasks, not CPU work. That points to disk/NFS, not CPU.
  • Rising 1-min vs 15-min average = the problem is getting worse.

🎯 Interview signal: “Linux load average includes uninterruptible I/O-wait tasks, so high load with idle CPUs means an I/O bottleneck — I’d look at iostat/iowait, not add CPU.” This one answer separates seniors from juniors.


3. CPU

top            # live: %us (user), %sy (system/kernel), %id (idle), %wa (I/O WAIT), %st (steal)
htop           # nicer top
mpstat -P ALL 1   # per-core utilization
pidstat 1      # per-process CPU over time

Reading top’s CPU line:

  • %us high → application/user code busy (real CPU work).
  • %sy high → lots of kernel/syscalls (context switches, I/O).
  • %wa high → CPU waiting on disk I/O (not a CPU problem — a storage one).
  • %st (steal) high → the hypervisor is giving your vCPU to other tenants (noisy neighbor / oversubscribed VM). ⚠️ On cloud VMs, high steal = you need a bigger/less-contended instance; it’s not your app.

4. Memory — and why free confuses people

free -h
#               total   used    free   shared  buff/cache   available
# Mem:           16Gi   4Gi     1Gi    0.2Gi   11Gi         11Gi

⚠️ “free” being low is normal and good. Linux uses spare RAM for buff/cache (disk cache) to speed things up, and reclaims it instantly when apps need memory. The number that matters is available — how much apps can actually get. “My server only has 1G free!” is usually fine if available is high.

vmstat 1        # si/so columns = swap in/out — nonzero = SWAPPING (bad)
cat /proc/meminfo
ps aux --sort=-%mem | head    # top memory consumers
  • Swapping (si/so nonzero, or high %wa) = out of RAM, pages going to disk → performance collapse (“thrashing”).

5. The OOM killer (guaranteed question)

When memory is truly exhausted (RAM + swap), the kernel’s Out-Of-Memory killer picks a process to kill to save the system. It scores processes (oom_score, roughly by memory use and oom_score_adj) and kills the highest.

dmesg | grep -i "killed process"      # OOM kill evidence
journalctl -k | grep -i oom
cat /proc/<pid>/oom_score

⚠️ “My process died with no error / exit 137” → often the OOM killer (or a SIGKILL; 137 = 128 + 9). In containers/Kubernetes, exceeding the memory limit triggers a cgroup OOM kill of that container (OOMKilled). The fix is right-sizing memory or fixing a leak, not just restarting.

🎯 Interview signal: “Exit code 137 = 128 + SIGKILL(9), classically the OOM killer or a cgroup memory-limit kill. I’d check dmesg/journalctl -k for the OOM message and the process’s oom_score.” Naming 137 = OOM is a strong signal.


6. Disk I/O

iostat -xz 1        # per-device: %util, await (latency), r/s w/s, aqu-sz (queue)
iotop               # per-process I/O (needs root)
sudo dstat          # combined cpu/disk/net/memory live view

Key columns in iostat -x:

  • %util near 100% → the disk is saturated.
  • await (ms) high → I/O latency is bad (slow disk, overloaded volume).
  • High disk I/O also shows as %wa in top and D-state processes.

⚠️ On cloud, EBS-type volumes have IOPS/throughput limits; hitting them looks like a slow disk even though the “disk” is fine — you’ve exhausted provisioned IOPS. Match iostat to the volume’s limits.


7. Network (perf angle)

ss -s               # socket summary
ip -s link          # interface errors/drops
sar -n DEV 1        # per-interface throughput over time
iftop / nload       # live bandwidth

High drops/errors on an interface, or saturated bandwidth, point to a network bottleneck (file 06 covers connectivity; this is throughput/saturation).


8. The “server is slow” playbook (say this)

1. uptime / load average → compare to nproc. Is load high? getting worse?
2. top → is it %us (CPU-bound app), %wa (I/O-bound), %sy (kernel), %st (noisy neighbor)?
3. free -h / vmstat → is 'available' low? any swapping (si/so)?
4. iostat -xz 1 → is a disk at 100% %util with high await?
5. Identify the culprit process: pidstat / iotop / ps --sort.
6. Correlate with logs (journalctl, app logs) and recent changes/deploys.
7. dmesg → OOM kills? hardware/driver errors?

🎯 Interview signal: a structured walk (resource by resource, top-down) with the reasoning at each step (“high %wa means I go to iostat, not add CPU”) is exactly what interviewers score.


Interview Questions

Q1. What does load average actually measure?

The average number of processes in the run queue — those running or runnable plus, on Linux, those in uninterruptible sleep (D state, usually waiting on disk/NFS I/O). It’s not a CPU percentage. You interpret it relative to core count: load equal to the number of cores is roughly fully utilized. Rising 1-minute vs 15-minute values mean the situation is worsening.

Q2. Load is very high but CPU usage is low. What’s happening?

The run queue is full of I/O-blocked (D-state) tasks, not CPU work — an I/O bottleneck. Because Linux includes uninterruptible-sleep tasks in load, a hung/slow disk or NFS mount inflates load while CPUs sit idle. I’d look at iostat -x for %util/await, top’s %wa, and which processes are in D state — not add CPU.

Q3. free -h shows almost no free memory. Is that a problem?

Usually not. Linux uses otherwise-idle RAM for buffer/page cache to speed up I/O and reclaims it instantly when applications need memory. The meaningful number is available, not free. A problem exists only if available is low and the system is swapping (nonzero si/so in vmstat), which indicates real memory pressure.

Q4. A process disappeared with exit code 137 and no application error. Why?

137 = 128 + 9, i.e., killed by SIGKILL — classically the kernel OOM killer reclaiming memory when the system (or a cgroup/container memory limit) ran out. I’d confirm with dmesg or journalctl -k for the “Killed process” / “Out of memory” message and check the process’s oom_score. Fix is right-sizing memory or fixing a leak, not just restarting.

Q5. Walk me through diagnosing “the server is slow.”

Start with uptime to see load vs core count and trend. Then top to classify: %us = CPU-bound app, %wa = I/O wait, %sy = kernel-heavy, %st = hypervisor steal on a VM. Check memory with free/vmstat (is available low, is it swapping?). If I/O-bound, iostat -xz 1 for %util/await, and iotop/pidstat to find the process. Correlate with logs, recent deploys, and dmesg for OOM/hardware errors. It’s resource-by-resource, top-down, with reasoning at each step.

Q6. What does high %st (steal) in top mean?

On a virtual machine, steal time is CPU the hypervisor gave to other tenants while your vCPU was ready to run — a sign of an oversubscribed host or a noisy neighbor. Your app isn’t the problem; the fix is a larger/less-contended instance or moving off the busy host.

Q7. How do you tell a disk-I/O bottleneck from a CPU bottleneck?

CPU-bound shows high %us with low %wa and the app on-CPU (pidstat). I/O-bound shows high %wa, D-state processes, and in iostat -x a device near 100% %util with high await. Load average can be high in both, so I distinguish them by whether time is spent on-CPU vs waiting on I/O.

Q8. (Senior) A cloud VM’s disk seems slow but the disk hardware is healthy. What do you suspect?

Provisioned IOPS/throughput exhaustion. Cloud block volumes (EBS-style) have per-volume IOPS and throughput limits (and burst credits on some tiers). If the workload exceeds them, iostat shows high await/%util even though nothing is “broken.” I’d compare observed IOPS/ throughput to the volume’s provisioned limits and either raise them, switch volume type, use a larger volume, or spread I/O — and check for burst-credit depletion on smaller volumes.


Next: 11 — Logging & Systematic Troubleshooting.