Scheduling & Automation

0%
15 mincronsystemd-timersautomation

Scheduling & Automation

Learn scheduling & automation concepts for DevOps interviews

13 — Scheduling & Automation

Recurring jobs — backups, cleanup, cert renewal, metrics — run via cron or systemd timers. Interviewers test cron syntax, the difference between cron and systemd timers, and the gotchas that make scheduled jobs silently fail.


1. cron — the classic scheduler

A cron daemon (crond/cron) runs commands on a schedule defined in crontabs.

crontab -e            # edit YOUR user crontab
crontab -l            # list your cron jobs
crontab -r            # remove (⚠️ deletes all — careful)
sudo crontab -e -u nginx   # edit another user's crontab

System-wide crontabs:

  • /etc/crontab and /etc/cron.d/* — system jobs (these have an extra user field).
  • /etc/cron.{hourly,daily,weekly,monthly}/ — drop a script in for periodic runs (run by run-parts).

The five-field syntax

┌─ minute (0–59)
│ ┌─ hour (0–23)
│ │ ┌─ day of month (1–31)
│ │ │ ┌─ month (1–12)
│ │ │ │ ┌─ day of week (0–7, 0 and 7 = Sunday)
│ │ │ │ │
* * * * *  command_to_run

Examples:

0 2 * * *        /opt/backup.sh              # every day at 02:00
*/5 * * * *      /opt/health-check.sh        # every 5 minutes
0 0 * * 0        /opt/weekly.sh              # Sundays at midnight
30 3 1 * *       /opt/monthly.sh             # 03:30 on the 1st of each month
0 9-17 * * 1-5   /opt/business-hours.sh      # top of each hour, 9–17, Mon–Fri
@reboot          /opt/on-boot.sh             # once at startup

🎯 Interview signal: fluently reading */5 * * * * (“every 5 minutes”) and 0 2 * * * (“daily at 2am”), and knowing /etc/crontab has an extra user field that user crontabs don’t.


2. Why cron jobs “work in the terminal but fail in cron” (the classic gotcha)

⚠️ cron runs with a minimal environment — a bare PATH (often just /usr/bin:/bin), no your shell profile, a different HOME, and no interactive niceties. So a script that works in your shell fails silently under cron.

Fixes / best practices:

  • Use absolute paths for every binary and file (/usr/bin/aws, not aws).
  • Set PATH (and needed env) at the top of the crontab or script.
  • Redirect output so you capture errors: >> /var/log/job.log 2>&1 (cron emails output by default, which usually goes nowhere on servers).
  • Log start/end and check exit codes.
# crontab with explicit environment + logging
PATH=/usr/local/bin:/usr/bin:/bin
MAILTO=""
0 2 * * * /opt/backup.sh >> /var/log/backup.log 2>&1

🎯 Interview signal: “The #1 reason a cron job fails is the stripped-down environment — no profile, bare PATH, different HOME. I use absolute paths and redirect stdout+stderr to a log.” This is the answer they’re fishing for.


3. at — run once at a future time

echo "/opt/deploy.sh" | at 22:00           # run once at 10pm today
at now + 1 hour                             # interactive
atq                                         # list pending
atrm <job>                                  # remove

Use at for one-off scheduling; cron/timers for recurring.


4. anacron — for machines that aren’t always on

cron assumes the machine is running at the scheduled time; if it’s off, the job is missed. anacron runs jobs that should happen periodically (daily/weekly) and catches up if the machine was off when they were due — used for laptops/desktops and the cron.daily etc. handlers on many distros.

⚠️ Interview nuance: “A server was off at 2am, did the daily backup run?” With plain cron, no — it’s simply skipped. anacron (or a systemd timer with Persistent=true) runs the missed job when the machine comes back.


5. systemd timers — the modern alternative

A timer unit (.timer) triggers a service unit (.service) on a schedule. Increasingly preferred over cron.

/etc/systemd/system/backup.service:

[Unit]
Description=Nightly backup
[Service]
Type=oneshot
ExecStart=/opt/backup.sh

/etc/systemd/system/backup.timer:

[Unit]
Description=Run backup daily at 2am
[Timer]
OnCalendar=*-*-* 02:00:00        # calendar schedule
Persistent=true                  # run missed jobs after downtime (like anacron)
RandomizedDelaySec=300           # jitter to avoid thundering-herd across a fleet
[Install]
WantedBy=timers.target
sudo systemctl enable --now backup.timer
systemctl list-timers                 # all timers + next/last run
journalctl -u backup.service          # the job's logs (built in!)

Why timers over cron (a strong interview answer)

cron systemd timer
Logging You redirect manually; else lost Automatic via journald (journalctl -u)
Missed runs (downtime) Skipped Persistent=true catches up
Dependencies/ordering None Full systemd deps (After=, Requires=)
Resource control None cgroup limits, Nice=, sandboxing
Randomized delay No RandomizedDelaySec (avoid fleet stampede)
Status/next run crontab -l only systemctl list-timers shows next/last

🎯 Interview signal: “systemd timers give me journald logging, Persistent=true for missed runs, dependency ordering, resource limits, and RandomizedDelaySec to avoid a fleet-wide thundering herd — cron gives none of that. I use timers for anything that matters.” ⚠️ Still know cron cold — it’s everywhere and simpler for quick jobs.


Interview Questions

Q1. Explain cron’s five fields with an example.

minute, hour, day-of-month, month, day-of-week, then the command. For example */5 * * * * runs every 5 minutes, 0 2 * * * runs daily at 02:00, and 0 0 * * 0 runs Sundays at midnight. /etc/crontab and /etc/cron.d add a user field before the command; personal crontabs (crontab -e) don’t, since they run as that user.

Q2. A script runs fine in your shell but fails as a cron job. Why?

cron runs with a minimal environment — a bare PATH, no login/profile, a different HOME, no interactive setup. So commands found via your shell’s PATH or relying on profile variables aren’t available. The fix is absolute paths for every binary/file, setting PATH and needed env in the crontab/script, and redirecting stdout+stderr to a log to capture errors.

Q3. The server was powered off at the scheduled cron time. Did the job run?

No — plain cron skips jobs whose scheduled time passed while the machine was off. To catch up missed runs you use anacron (for daily/weekly periodic jobs) or a systemd timer with Persistent=true, which runs the job when the system comes back.

Q4. cron vs systemd timers — when and why?

systemd timers add automatic journald logging (journalctl -u), Persistent=true to run missed jobs after downtime, dependency ordering and resource/sandbox controls, and RandomizedDelaySec to stagger runs across a fleet. cron is simpler and ubiquitous. I use timers for anything that needs logging, reliability, or ordering, and cron for quick, simple periodic tasks — but I know both.

Q5. How do you schedule a one-time job in the future?

at — e.g., echo "/opt/job.sh" | at 22:00. atq lists pending jobs and atrm removes them. cron/timers are for recurring schedules; at is for one-offs.

Q6. How do you capture the output/errors of a scheduled job?

For cron, redirect explicitly: >> /var/log/job.log 2>&1 (cron otherwise emails output, which usually goes nowhere on servers). For systemd timers it’s automatic — the service’s stdout/stderr are captured by journald and viewable with journalctl -u <service>. I also log start/end and check exit codes so failures are visible.

Q7. How do you avoid a fleet of servers all hammering a backend at exactly 2am?

Add jitter. systemd timers support RandomizedDelaySec to spread executions over a window; with cron you’d add a random sleep at the top of the job. This prevents a thundering-herd load spike on shared backends (DB, artifact store) when many hosts run the same job on the same schedule.

Q8. (Senior) Design a reliable nightly backup job. What do you use and why?

A systemd timer + oneshot service: OnCalendar for the schedule, Persistent=true so a missed run (host down) still executes on boot, RandomizedDelaySec to avoid fleet stampede, and journald for built-in logging (journalctl -u backup). The script itself is hardened (set -euo pipefail, absolute paths, flock so runs don’t overlap, exit-code checks, alert-on-failure). I’d add monitoring that alerts if the backup didn’t succeed within its window and verify restores periodically — an untested backup isn’t a backup.


Next: 14 — Containers & Linux Internals — how Docker and Kubernetes are “just Linux.”