Incident Overview
A platform team declares a fictional SEV-2 incident after its central logging view reports a sharp fall in accepted events. The incident commander sees a dashboard indicating that incoming volume has dropped, while service owners report that customer-facing requests continue to complete normally.
A separate delivery counter shows events entering the collection boundary, but the searchable-event counter remains behind.The commander must choose one immediate action. The operational assumptions are explicit: the team is working in an isolated validation environment; product version and permissions have not yet been confirmed; no configuration change is authorised; and the displayed signals may cover different stages or time windows.
The objective is to preserve evidence, define the affected boundary and decide whether recovery or escalation is justified.
Investigation Options
Review the available operational moves and select the best immediate action.
Freeze logging changes, record the dashboard scope and compare matching time windows at the collection and search boundaries.
Escalate immediately as total customer-service failure based only on the central volume alert.
Declare the alert harmless because service owners report successful requests.
Restart a logging component before confirming its version, ownership or failure boundary.