Incident scenarios
Guided troubleshooting walkthroughs modelled on incidents that actually happen on call. Each step asks what you would check next, then explains the reasoning and shows the commands. Ends with root cause, fix and prevention.
6 scenarios · 3 free
6 scenarios
- 6 steps →AWSMediumFree~12 min
ALB returning 502 intermittently
Roughly 15% of requests to an Application Load Balancer return 502 Bad Gateway, but only for one of three target groups. Health checks pass.
- 5 steps →LinuxMediumFree~10 min
Instance disk full despite low reported usage
A production instance reports 'No space left on device' while `df -h` shows 60% free space. Writes are failing on one service only.
- 5 steps →KubernetesMediumPremium~14 min
Pods stuck in CrashLoopBackOff
After a routine image tag bump, roughly half the pods in a namespace are in CrashLoopBackOff. Events show the container starts and exits with code 1 within seconds.
- 4 steps →CI/CDMediumPremium~11 min
CI pipeline red intermittently for the same commit
A pipeline that passes 9 times out of 10 fails intermittently on the same commit with a test timeout, yet the commit works locally every time.
- 4 steps →NetworkingMediumFree~8 min
Container cannot reach the database
An application container reports `ECONNREFUSED` connecting to a Postgres database that is running on the same Docker host and is reachable from the host itself.
- 5 steps →NetworkingHardPremium~15 min
Intermittent TLS handshake failures
About 3% of HTTPS requests from a mobile client fail with `SSL handshake failure`. The same client talking to other hosts is fine, and failures cluster in the morning.