Home/Linux Administration/Troubleshooting
🧭

Level 21 of 24

Troubleshooting

Symptom → cause → command → fix → verify → prevent, for the failures you’ll actually hit.

This level is a synthesis of everything before it, organized around real failure scenarios rather than individual tools — the goal is a repeatable framework, applied to the specific failures you'll actually encounter.

Prerequisites

  • • Levels 1–20

By the end of this level, you can

  • ✓ Apply a consistent symptom-to-resolution framework to any failure
  • ✓ Diagnose the most common real-world Linux failure scenarios specifically
  • ✓ Know which command to reach for first for a given category of symptom

A systematic troubleshooting framework

Symptom → possible causes → commands → interpretation → resolution → verification → prevention. · 8 min

Effective troubleshooting is a discipline, not an instinct — the difference between an administrator who resolves an incident in ten minutes and one who takes two hours is usually the presence or absence of a systematic approach, not raw knowledge. The framework: observe the actual symptom precisely (not "the site is down" but "requests to /api return a 502 starting at 14:03"), list plausible causes before investigating any single one (resist tunnel vision on the first theory), check the most likely and cheapest-to-check causes first, interpret command output rather than pattern-matching on familiar-looking text, apply the narrowest fix that addresses the actual root cause, verify the fix actually resolved the symptom (not just that the fix command "ran successfully"), and finally consider what would prevent recurrence.

The single most valuable habit inside this framework: narrow by time and by layer before guessing. Time — when exactly did this start, and what changed around then (a deploy, a config change, a scheduled job)? Layer — is this a network problem, an application problem, a permissions problem, a resource-exhaustion problem? Every scenario in the rest of this level follows this shape explicitly.

Resist the pressure to apply a fix before you've confirmed the actual cause, even under incident pressure — restarting a service "just in case" can temporarily mask a symptom while leaving the actual root cause (a memory leak, a full disk, a bad deploy) completely unaddressed and guaranteed to recur, often at a worse moment.

The framework, compressed

Symptom (precise, timestamped)
  ↓
Possible causes (list several, don't fixate on the first)
  ↓
Commands (check cheapest/most-likely first)
  ↓
Interpretation (read the actual output, don't pattern-match)
  ↓
Resolution (the narrowest fix for the actual root cause)
  ↓
Verification (confirm the symptom is actually gone)
  ↓
Prevention (what stops this from recurring?)

Takeaway: Narrow by time (what changed, and when) and by layer (network, application, permissions, resources) before guessing at a fix — and always verify a fix actually resolved the symptom rather than just having "run successfully".

Common real-world failure scenarios

SSH unreachable, disk full, service failed, DNS broken, and permission denied — worked through explicitly. · 14 min

SSH unreachable: check in order — is the server up at all (ping, or the cloud provider's console)? Is sshd actually running (`systemctl status sshd` via the console/other access if SSH itself is down)? Is a firewall (OS-level or cloud Security Group/NSG) blocking port 22? Did a recent `sshd_config` change (Level 9) introduce a syntax error or overly restrictive setting? A cloud provider's browser-based console access (bypassing SSH entirely) is often the only way back in if SSH itself is genuinely broken.

Disk full: `df -h` identifies which filesystem is full; `du -sh /path/*` narrows down which directory is the actual culprit — `/var/log` and unmanaged application data directories are common offenders (Level 12). The immediate fix is freeing space (clearing old logs, rotating manually if logrotate wasn't configured); the actual root-cause fix is making sure that growth pattern doesn't recur unmanaged.

Service failed: `systemctl status <service>` first, `journalctl -u <service> --since "recent"` second (Level 12) — read the actual error rather than just restarting blindly. A service that fails immediately after a config change is almost always the change itself (test config syntax before applying, as covered for both sshd and Nginx); a service that fails after running fine for a while more often points to resource exhaustion or an external dependency (database, disk) that became unavailable.

DNS failure: confirm whether it's actually DNS by testing with a raw IP directly (`curl -I https://93.184.216.34` style bypass, or `dig`/`nslookup` against a known-working resolver like 8.8.8.8 versus the system's configured resolver) — this isolates whether the problem is DNS resolution specifically or broader connectivity. Check `/etc/resolv.conf` for a genuinely misconfigured or unreachable nameserver.

Permission denied: this course's own Level 4 and Level 11 apply directly — check standard permissions and ownership first (`ls -l`), then, on RHEL-family systems specifically, check SELinux context and `ausearch` if standard permissions look correct but access is still denied. A "permission denied" that persists despite seemingly correct standard permissions, on a RHEL-family system, is the classic SELinux signature covered in Level 11.

  • • SSH down → check server up, sshd running, firewall/NSG, sshd_config syntax — use cloud console as fallback access
  • • Disk full → df -h to find the filesystem, du -sh to find the actual culprit directory
  • • Service failed → systemctl status first, journalctl -u --since second, read the actual error
  • • DNS broken → test with a raw IP or a known-good resolver to isolate DNS specifically from general connectivity
  • • Permission denied → standard permissions/ownership first, then SELinux context on RHEL-family systems

Quick Check

A service fails to start immediately after you edited its configuration file. What's the most likely category of cause?

Takeaway: For each of these scenarios, the diagnostic command comes before any fix — restarting a service, or making a firewall change, before reading what actually failed usually wastes a cycle or masks the real cause.

Sponsor / Advertisement