Incident Overview
At 09:14 on a fictional Monday, the monitoring platform for a mid-sized retail application fires two alerts within ninety seconds of each other for the same host, app-node-3. The first alert states Load Average Critical: 22.4 (1m).
The second alert, from a separate CPU utilisation check, states CPU Utilisation Normal: 6%. The on-call graduate administrator, Tomasz, is asked to triage the host before deciding whether to restart any services or escalate to the platform team.
The dashboards appear to disagree with each other, and Tomasz has thirty minutes before the next deployment window opens on an adjacent host.
Investigation Options
Review the available operational moves and select the best immediate action.
Immediately restart the application service on app-node-3 to clear the D-state processes
Collect and correlate uptime, mpstat, ps and iostat evidence before deciding on any remediation
Silence the load average alert as a false positive and continue monitoring CPU utilisation only
Immediately reduce the application's logging verbosity in production without reviewing current disk usage