Systemd & Services

0%
20 minsystemdservicesinit

Systemd & Services

Learn systemd & services concepts for DevOps interviews

04 — systemd & Service Management

systemd is where most “my service won’t start / won’t stay up / won’t come back after reboot” interview questions live. This file makes you fluent in units, dependencies, restart policy, journalctl, and timers — the daily bread of Linux ops.


1. What systemd manages: units

systemd models everything as a unit. The type is the file suffix:

Unit type Purpose
.service A daemon/process to run (nginx, sshd)
.socket A socket that can start a service on demand
.target A group of units / a system state (file 00)
.timer A scheduled trigger (cron replacement, file 13)
.mount / .automount Filesystem mounts
.path Watch a path and trigger a unit on change
.device Kernel devices exposed as units
systemctl list-units --type=service          # active services
systemctl list-unit-files --state=enabled     # what starts at boot

2. The commands you must know cold

systemctl start   nginx        # start now
systemctl stop    nginx        # stop now
systemctl restart nginx        # stop + start
systemctl reload  nginx        # re-read config WITHOUT dropping connections (if supported)
systemctl status  nginx        # state + recent log lines + PID + cgroup
systemctl enable  nginx        # start automatically at boot (creates a symlink)
systemctl disable nginx        # don't start at boot
systemctl enable --now nginx   # enable AND start in one shot
systemctl is-active nginx      # active/inactive
systemctl is-enabled nginx     # enabled/disabled
systemctl daemon-reload        # RELOAD systemd itself after editing unit files
systemctl cat nginx            # show the effective unit file

⚠️ Two distinctions interviewers probe:

  • start ≠ enable. start runs it now; enable makes it start at boot. You usually want enable --now. “Service is running but didn’t come back after reboot” = it was started but never enabled.
  • reload ≠ restart. reload re-reads config without killing the process (no downtime, if the daemon supports it); restart fully stops and starts (brief downtime, resets connections).
  • After editing a unit file you must run systemctl daemon-reload or systemd uses the old definition. Forgetting this is a classic “my changes did nothing.”

🎯 Interview signal: “enable --now and daemon-reload are the two things people forget — enable for persistence across reboots, daemon-reload after editing units.”


3. Anatomy of a service unit

/etc/systemd/system/myapp.service:

[Unit]
Description=My App API
After=network-online.target          # ordering: start after the network is up
Wants=network-online.target          # weak dependency (don't fail if it's missing)
Requires=postgresql.service          # strong dependency (fail if this fails)

[Service]
Type=simple                          # main process runs in foreground (most common)
User=myapp                           # run as an unprivileged user (least privilege)
Group=myapp
WorkingDirectory=/opt/myapp
ExecStart=/opt/myapp/bin/api --port 8080
ExecReload=/bin/kill -HUP $MAINPID
Restart=on-failure                   # auto-restart if it crashes
RestartSec=5s                        # wait 5s between restarts
Environment=LOG_LEVEL=info
EnvironmentFile=/etc/myapp/env       # load env from a file
# Hardening (file 12):
NoNewPrivileges=true
ProtectSystem=strict
PrivateTmp=true

[Install]
WantedBy=multi-user.target           # enable => start at multi-user (normal server boot)

Key fields:

  • After= / Before= = ordering only (when). Requires= / Wants= = dependency (whether it’s pulled in / must succeed). ⚠️ After does not imply Requires — a common confusion. You often want both.
  • Type=: simple (foreground), forking (daemon that double-forks), oneshot (runs and exits, e.g., a setup task), notify (tells systemd when ready), exec.
  • WantedBy=multi-user.target in [Install] is what makes enable work — it creates the symlink that pulls the service in at boot.

Placement:

  • /etc/systemd/system/ — your custom/overriding units (highest precedence).
  • /usr/lib/systemd/system/ — units shipped by packages (don’t edit; override instead).
systemctl edit myapp     # create a drop-in override (safer than editing the shipped unit)

4. Restart policy & start-rate limiting (a favorite scenario)

Restart=on-failure       # also: no, always, on-abnormal, on-success
RestartSec=5s
StartLimitIntervalSec=60
StartLimitBurst=5        # if it restarts >5 times in 60s, systemd gives up

⚠️ If a service keeps crashing, systemd will restart it a few times, then rate-limit and mark it failed with “start request repeated too quickly.” The service then won’t come back until you fix the cause and run systemctl reset-failed <unit>. Interviewers love: “the service is dead and won’t restart even though Restart=always” → hit the start limit.

systemctl reset-failed myapp    # clear the failed/rate-limited state after fixing

🎯 Interview signal: explaining that Restart= + StartLimitBurst interact — auto-restart is bounded so a crash-loop doesn’t hammer the box forever — shows real systemd depth.


5. journald — logs are built in

systemd captures each service’s stdout/stderr into the journal (file 11 goes deeper).

journalctl -u myapp                 # all logs for a unit
journalctl -u myapp -f              # follow (like tail -f)
journalctl -u myapp --since "10 min ago"
journalctl -u myapp -p err          # only error priority and above
journalctl -b                       # current boot; -b -1 = previous boot
journalctl -u myapp -n 50 --no-pager

⚠️ By default the journal may be volatile (lost on reboot) unless /var/log/journal exists (persistent storage). Set Storage=persistent in /etc/systemd/journald.conf to keep logs across reboots — important for post-mortems.


6. Debugging a service that won’t start (say this method)

systemctl status myapp          # 1. state + last log lines + exit code
journalctl -u myapp -n 100      # 2. full recent logs — the real error
systemctl cat myapp             # 3. verify the effective unit (ExecStart path, User)
sudo -u myapp /opt/myapp/bin/api  # 4. run the command manually AS the service user
systemctl daemon-reload         # 5. if you edited the unit and forgot this

Common root causes: wrong ExecStart path, missing permissions for User=, missing EnvironmentFile, a dependency (DB) not up, or a 203/EXEC error (binary not executable / wrong path) or 217/USER (user doesn’t exist).

🎯 Interview signal: “203/EXEC means systemd couldn’t execute the binary — usually a bad path or the file isn’t executable; 217/USER means the User= doesn’t exist.” Naming exit codes shows hands-on experience.


7. Targets and dependencies (recap + control)

systemctl get-default                    # default boot target
systemctl list-dependencies multi-user.target
systemctl isolate rescue.target          # switch state now
systemctl mask  bluetooth.service        # fully disable (symlink to /dev/null) — can't be started
systemctl unmask bluetooth.service

⚠️ mask is stronger than disable — a masked unit can’t be started at all, even as a dependency. Useful to guarantee something never runs.


Interview Questions

Q1. What’s the difference between systemctl start and systemctl enable?

start runs the service now; enable configures it to start automatically at boot (by creating the WantedBy symlink). They’re independent — a service can be running now but not enabled (won’t survive reboot) or enabled but not currently running. Usually you want enable --now to do both.

Q2. Difference between reload and restart?

reload tells the service to re-read its configuration without stopping — no dropped connections — but only works if the daemon implements it (via ExecReload, often a SIGHUP). restart fully stops and starts the process, causing brief downtime and resetting connections. Prefer reload for config changes when supported.

Q3. You edited a unit file but your change had no effect. Why?

systemd caches unit definitions; after editing a unit file you must run systemctl daemon-reload so it re-reads them, then restart the service. Also, if you edited the package-shipped unit in /usr/lib/systemd/system a drop-in override in /etc/systemd/system may be overriding it — check systemctl cat.

Q4. Difference between After= and Requires=?

After=/Before= control ordering — when a unit starts relative to another — but don’t create a dependency. Requires=/Wants= create a dependency — Requires is strong (if the dependency fails, this unit fails) and Wants is weak (start it if possible, don’t fail otherwise). They’re orthogonal: you often set both Requires=db.service and After=db.service so it starts after and depends on the DB.

Q5. A service with Restart=always is dead and won’t come back. Explain.

It hit the start rate limit. systemd restarts a crashing service up to StartLimitBurst times within StartLimitIntervalSec; exceeding that marks it failed with “start request repeated too quickly” and stops restarting to avoid a crash-loop hammering the box. After fixing the root cause, clear it with systemctl reset-failed <unit>.

Q6. How do you view logs for a specific service, including only errors since the last boot?

journalctl -u <unit> -b -p err. -u scopes to the unit, -b to the current boot, -p err to error priority and above. Add -f to follow live or --since for a time window.

Q7. Walk me through debugging a service that fails to start.

systemctl status for state, exit code, and last log lines; journalctl -u <unit> -n 100 for the real error; systemctl cat to verify ExecStart path and User; then run the ExecStart command manually as the service user to reproduce. Common causes: bad binary path or non-executable (203/EXEC), nonexistent User= (217/USER), missing EnvironmentFile, permissions, or a dependency not up. Remember daemon-reload if the unit was edited.

Q8. What’s the difference between disable and mask?

disable stops a unit from starting at boot but it can still be started manually or pulled in as a dependency. mask symlinks the unit to /dev/null so it can’t be started at all, even as a dependency — a hard guarantee that it never runs. Unmask to reverse it.

Q9. (Senior) Why run a service as a dedicated non-root user with ProtectSystem/PrivateTmp, and how?

Least privilege: if the service is compromised, the blast radius is limited to that unprivileged account and a sandboxed view of the filesystem. In the unit I set User=/ Group= to a system account, NoNewPrivileges=true (can’t gain privileges via setuid), ProtectSystem=strict (read-only OS dirs), PrivateTmp=true (isolated /tmp), and can add ReadWritePaths= for the few dirs it legitimately needs. systemd enforces these via namespaces/cgroups, so hardening is declarative in the unit file.


Next: 05 — Package Management — installing and managing software the distro way.