A systematic troubleshooting framework
Symptom → possible causes → commands → interpretation → resolution → verification → prevention. · 8 min
Effective troubleshooting is a discipline, not an instinct — the difference between an administrator who resolves an incident in ten minutes and one who takes two hours is usually the presence or absence of a systematic approach, not raw knowledge. The framework: observe the actual symptom precisely (not "the site is down" but "requests to /api return a 502 starting at 14:03"), list plausible causes before investigating any single one (resist tunnel vision on the first theory), check the most likely and cheapest-to-check causes first, interpret command output rather than pattern-matching on familiar-looking text, apply the narrowest fix that addresses the actual root cause, verify the fix actually resolved the symptom (not just that the fix command "ran successfully"), and finally consider what would prevent recurrence.
The single most valuable habit inside this framework: narrow by time and by layer before guessing. Time — when exactly did this start, and what changed around then (a deploy, a config change, a scheduled job)? Layer — is this a network problem, an application problem, a permissions problem, a resource-exhaustion problem? Every scenario in the rest of this level follows this shape explicitly.
Resist the pressure to apply a fix before you've confirmed the actual cause, even under incident pressure — restarting a service "just in case" can temporarily mask a symptom while leaving the actual root cause (a memory leak, a full disk, a bad deploy) completely unaddressed and guaranteed to recur, often at a worse moment.
The framework, compressed
Symptom (precise, timestamped)
↓
Possible causes (list several, don't fixate on the first)
↓
Commands (check cheapest/most-likely first)
↓
Interpretation (read the actual output, don't pattern-match)
↓
Resolution (the narrowest fix for the actual root cause)
↓
Verification (confirm the symptom is actually gone)
↓
Prevention (what stops this from recurring?)Takeaway: Narrow by time (what changed, and when) and by layer (network, application, permissions, resources) before guessing at a fix — and always verify a fix actually resolved the symptom rather than just having "run successfully".