Logging & Troubleshooting
Logging & Troubleshooting
Learn logging & troubleshooting concepts for DevOps interviews
11 — Logging & Systematic Troubleshooting
Logs are your evidence, and a repeatable troubleshooting method is what interviewers really test with “how would you debug X?” This file covers
journalctl,/var/log, log rotation, and a debugging methodology you can apply to any incident.
1. Two logging worlds: journald and /var/log
Modern Linux has both:
- systemd-journald — structured, binary journal capturing each service’s stdout/stderr
plus metadata (unit, PID, priority). Query with
journalctl. /var/log/*— traditional plain-text logs, written by apps directly or by a syslog daemon (rsyslog/syslog-ng). journald often forwards to rsyslog too.
Service stdout/stderr ─► journald (binary, queryable)
│ (optionally forwards)
▼
rsyslog ─► /var/log/{syslog,messages,auth.log,...}
🎯 Interview signal: “journald is the structured, per-unit journal; /var/log is the
traditional text logs. Many apps (nginx, app servers) still write their own files under
/var/log, so I check both.” Knowing both worlds coexist is the mature answer.
2. journalctl — query the journal
journalctl -u nginx # logs for a unit
journalctl -u nginx -f # follow live
journalctl -u nginx --since "1 hour ago" --until "10 min ago"
journalctl -u nginx -p err # priority err and above (emerg..debug)
journalctl -b # this boot; -b -1 = previous boot
journalctl -k # kernel messages (like dmesg)
journalctl -p err -b # all errors this boot
journalctl _PID=1234 # by field (structured!)
journalctl -u nginx -o json-pretty # structured output
journalctl --disk-usage # how big the journal is
Priority levels (-p): emerg(0) alert(1) crit(2) err(3) warning(4) notice(5) info(6) debug(7).
⚠️ By default the journal may be volatile (RAM, lost on reboot). Create
/var/log/journal or set Storage=persistent in /etc/systemd/journald.conf for logs that
survive reboots — essential for post-mortems.
3. Key files in /var/log
| File (Debian / RHEL) | Contains |
|---|---|
/var/log/syslog / /var/log/messages |
General system messages |
/var/log/auth.log / /var/log/secure |
Authentication, sudo, SSH logins ← security |
/var/log/kern.log |
Kernel messages |
/var/log/dmesg |
Boot-time kernel ring buffer |
/var/log/nginx/, /var/log/mysql/ |
Per-app logs |
tail -f /var/log/syslog
grep "Failed password" /var/log/auth.log # brute-force attempts (file 12)
dmesg -T | tail # kernel ring buffer with timestamps
🎯 auth.log/secure is where you spot SSH brute-force (“Failed password” floods) and sudo
misuse — a common security-troubleshooting question.
4. Log rotation — stop logs from filling the disk
Unbounded logs are the #1 cause of “disk full” (file 07). logrotate rotates, compresses, and deletes old logs on a schedule (run via cron/systemd timer).
/etc/logrotate.d/myapp:
/var/log/myapp/*.log {
daily
rotate 14 # keep 14 days
compress # gzip old logs
delaycompress
missingok
notifempty
copytruncate # OR use postrotate to signal the app to reopen its log
}
⚠️ The deleted-but-open trap (file 01/07) applies: if logrotate just moves/deletes the log but the app keeps writing to the old file descriptor, space isn’t freed and new logs go nowhere. Solutions:
copytruncate: copy then truncate the original in place (small race window), orpostrotate: send the app a signal (oftenSIGHUPorUSR1) to reopen its log file. journald-managed services avoid this since journald owns the stream.
🎯 Interview signal: explaining copytruncate vs postrotate-signal-to-reopen shows you
understand why naive rotation loses logs.
5. A repeatable troubleshooting methodology (the real interview answer)
When asked “how would you debug X,” don’t jump to a command — show a method:
1. DEFINE the problem precisely.
- What's the symptom? Since when? What changed (deploy, config, patch, traffic)?
- Is it total or intermittent? Affecting one host or many?
2. REPRODUCE / confirm the scope.
- Can you trigger it? Which users/hosts/requests are affected?
3. GATHER evidence — work the layers, don't guess.
- Logs: journalctl -u <svc> -p err ; app logs ; auth.log ; dmesg
- State: systemctl status ; ps ; ss -tlnp ; df -h/-i ; free -h ; top ; iostat
4. FORM a hypothesis from evidence, then TEST it (change one thing).
5. FIX, then VERIFY the symptom is gone and no new issue appeared.
6. DOCUMENT: root cause, fix, and prevention (monitoring/alert/runbook).
Guiding principles interviewers love to hear:
- “What changed?” is the highest-yield first question — most incidents follow a change.
- Bisect / narrow the scope: is it DNS or network (file 06)? app or infra? one host or fleet? Halve the problem space each step.
- Change one variable at a time so you know what fixed it.
- Correlate timelines: match the symptom’s start to deploys, cron jobs, traffic spikes, or log timestamps.
🎯 Interview signal: leading with “first, what changed and when did it start?” and then a layered evidence-gathering plan — instead of a random command — is exactly the SRE mindset they’re screening for.
6. Worked example: “the website returns 502”
1. Scope: all users? one endpoint? since a deploy 10 min ago? → "since deploy" is a big clue.
2. Front to back:
- LB/proxy logs (nginx): 502 = upstream (app) unreachable/erroring.
- Is the app service up? systemctl status myapp ; journalctl -u myapp -p err
- Is it listening? ss -tlnp | grep 8080
- Resource pressure? free -h (OOM? dmesg 137) ; df -h (disk full breaking writes?)
3. Hypothesis: app crashed on the new deploy (bad config / OOM).
4. Test: roll back the deploy or fix config; watch journalctl.
5. Verify: 502s stop, health check green.
6. Document + add an alert on app-down and upstream 5xx.
⚠️ 502/504 from a proxy almost always means the backend is down/slow/unreachable — check the app, not the proxy first (connection refused = app down; timeout = app slow or firewall, per file 06).
Interview Questions
Q1. Where do you look for logs on a modern Linux system?
Two places: the systemd journal via
journalctl(structured, per-unit, captures services’ stdout/stderr with metadata) and traditional text files under/var/log(syslog/messages, auth.log/secure, kernel, and per-app logs like nginx). Many apps still write their own/var/logfiles, and journald often forwards to rsyslog, so I check both.
Q2. How do you see only errors for a service since the last boot, live?
journalctl -u <unit> -b -p err -f:-uscopes to the unit,-bto the current boot,-p errto error priority and above, and-ffollows new entries. For a window I’d use--since/--until.
Q3. Why might journal logs disappear after a reboot, and how do you keep them?
By default the journal can be volatile (stored in RAM) if
/var/log/journaldoesn’t exist, so it’s lost on reboot. Create that directory or setStorage=persistentin journald.conf to retain logs across reboots — important for post-mortems.
Q4. How does log rotation work and what’s the classic pitfall?
logrotate (via cron/timer) rotates, compresses, and prunes logs per rules (daily/size, keep N, compress). The pitfall is the deleted-but-open problem: if it just renames/removes the file while the app holds the old fd, disk space isn’t freed and new logs go nowhere. Fixes are
copytruncate(copy then truncate in place) or apostrotatesignal telling the app to reopen its log file.
Q5. Walk me through your general troubleshooting method.
Define the symptom precisely and ask what changed and when — most incidents follow a deploy/config/patch. Confirm the scope (total vs intermittent, one host vs fleet). Gather evidence layer by layer — logs (journalctl, app logs, auth.log, dmesg) and state (systemctl, ss, df, free, top, iostat) — rather than guessing. Form a hypothesis from that evidence, test by changing one thing, fix, verify the symptom is gone, then document root cause and add prevention (alert/runbook).
Q6. A proxy returns 502/504. Where do you look first?
The backend, not the proxy. 502/504 means the upstream app is unreachable, erroring, or too slow. I’d check whether the app service is up and listening (
systemctl status,ss -tlnp), its error logs (journalctl -u app -p err), and resource pressure (OOM in dmesg, disk full). A connection refused points to the app being down; a timeout points to it being slow or a firewall — and I’d correlate with any recent deploy.
Q7. Where do you check for SSH brute-force or sudo misuse?
/var/log/auth.logon Debian or/var/log/secureon RHEL (also viajournalctl). Floods of “Failed password” entries indicate brute-force attempts; sudo invocations and authentication events are logged there too. That feeds into hardening — key-only SSH, fail2ban (file 12).
Q8. (Senior) An intermittent issue happens a few times a day with no obvious trigger. How do you catch it?
Since it’s not reproducible on demand, I instrument to capture it when it happens: ensure persistent journald, raise relevant log verbosity temporarily, add metrics/alerts around the suspected subsystem, and correlate timestamps across app logs, system metrics, and any scheduled jobs (cron/timers) or traffic patterns. I’d look for periodicity (a cron job, a backup, a cache expiry, a neighbor’s load) and, if needed, capture state on trigger (a script that dumps
ss,top,dmesg, thread dumps when an error signature appears). The goal is to turn “intermittent and mysterious” into “captured with evidence,” then bisect from there.
Next: 12 — Security & Hardening.