Home/Linux Administration/Logging & Monitoring
📊

Level 12 of 24

Logging & Monitoring

journald, rsyslog, log rotation, and reading logs to find the actual root cause.

When something breaks, the logs are almost always where the actual answer is sitting — the skill is knowing which log, which command, and which time window to look at, rather than scrolling blindly.

Prerequisites

  • • Level 5: Processes & Services

By the end of this level, you can

  • ✓ Use journalctl and /var/log effectively to find the actual cause of a failure
  • ✓ Explain what logrotate does and why unmanaged logs are a real production risk
  • ✓ Correlate a service failure with its logs using a specific time window

journald, rsyslog, and /var/log

Where logs actually live, and the fastest path to the relevant line. · 10 min

On modern systemd-based systems, `journald` captures logs centrally in a binary format, queried through `journalctl` — this includes systemd's own messages, kernel messages, and anything a service run under systemd writes to standard output or error. Many traditional services also still write plain-text logs directly to files under `/var/log` (`/var/log/nginx/error.log`, `/var/log/auth.log` or `/var/log/secure` for authentication attempts) — some distributions configure `rsyslog` to also mirror journald content into these traditional flat files for compatibility with older tooling and habits.

The fastest, most targeted path to a relevant log line: `journalctl -u <service>` scopes to one service instead of the entire system journal. `journalctl --since "10 minutes ago"` (or `--since "2026-09-19 14:00" --until "2026-09-19 14:15"`) scopes by time — essential when you know roughly when an incident happened and don't want to wade through unrelated noise. `journalctl -p err` filters to error-priority-and-above messages only. `journalctl -f` follows new entries live, the systemd equivalent of `tail -f`.

For plain-text logs under `/var/log`, `tail -f /var/log/nginx/error.log` follows a specific file live, and `grep` narrows a large log file to lines matching a pattern — `grep -i "error" /var/log/syslog | tail -50` is a common, genuinely useful combination for a first pass over a noisy log.

CommandPurposeExample
journalctl -u nameView logs for a specific systemd service—
journalctl -fFollow the journal live—
journalctl --since "1 hour ago"Scope logs to a relative or absolute time window—
journalctl -p errShow only error-priority-and-above entries—
tail -f /var/log/fileFollow a plain-text log file live—
grep -i pattern fileSearch a log file, case-insensitivegrep -i "failed" /var/log/auth.log
🧪

Hands-On Lab

Troubleshoot a failed service using logs alone

Objectives

  • ✓ Intentionally break a service's configuration
  • ✓ Diagnose the cause purely from journalctl output, without guessing
  • ✓ Fix it and confirm via the logs that it's resolved

Instructions

  1. Install nginx (`sudo apt install nginx`) and confirm it's running with `systemctl status nginx`.
  2. Introduce a deliberate syntax error into `/etc/nginx/nginx.conf` (e.g. remove a closing brace) and attempt `sudo systemctl restart nginx`.
  3. Notice the restart fails — run `sudo journalctl -u nginx --since "5 minutes ago"` and read the actual error nginx reported, rather than guessing.
  4. Fix the syntax error based on what the log actually said, then restart and confirm success both via `systemctl status` and by seeing no new errors in `journalctl -u nginx`.
Hints (1)
  • nginx also has its own dedicated error log at `/var/log/nginx/error.log`, which can carry more detail than the systemd journal alone — check both when journalctl output seems incomplete.

Takeaway: `journalctl -u <service> --since "recent time window"` is the fastest path to the actual cause of a service failure — narrow by service and time before reading anything, instead of scrolling through the entire system log.

Log rotation and why unmanaged logs are a real risk

logrotate, and the specific failure mode of a disk filling up from logs alone. · 6 min

Logs grow continuously as long as a service runs, and without active management, a busy service's log file can grow to consume all available disk space — a genuinely common real-world outage cause, distinct from any application bug: the disk fills, and every service on that filesystem starts failing writes, often with confusing, unrelated-looking error messages that don't obviously point back to "the disk is full".

`logrotate` is the standard tool that solves this: configured per-service (commonly in `/etc/logrotate.d/`), it periodically renames the current log file, optionally compresses the rotated-out version, deletes logs older than a configured retention period, and signals the service to start writing to a fresh file — all automatically, on a schedule, with no manual intervention once configured correctly.

When troubleshooting "the disk is full" as a root cause (not just a symptom), `du -sh /var/log/*` quickly shows which specific log is the actual culprit, and checking whether that service has a logrotate configuration at all — or one with a retention/size setting that's simply too generous for its actual log volume — is usually the real, permanent fix, rather than a one-time manual cleanup that will recur.

CommandPurposeExample
du -sh /var/log/*Show which log files/directories are consuming the most space—
logrotate -d /etc/logrotate.confDry-run logrotate's configuration to see what it would do, without actually doing it—
df -hConfirm actual current disk usage across mounted filesystems—

Takeaway: A disk filling up from unmanaged logs is a real, common outage cause, not a hypothetical — check that every noisy service has a sensible logrotate configuration before it becomes the reason a server goes down.

Sponsor / Advertisement