Production Incident Triage
Concise breakdowns of real-world production incidents. We explore the diagnostic evidence, the root causes, and the deterministic remediation steps taken to restore service.
Latest Briefing
Kubernetes HPA Scales Pods Up While CPU Graphs Show Idle Capacity
A fictional Kubernetes triage exercise in which a Horizontal Pod Autoscaler keeps scaling a checkout service up even though CPU dashboards show idle capacity, revealing a stale custom-metric scrape mismatch.
Incident Archive
30 RECORDSKubernetes HPA Scales Pods Up While CPU Graphs Show Idle Capacity
A fictional Kubernetes triage exercise in which a Horizontal Pod Autoscaler keeps scaling a checkout service up even though CPU dashboards show idle capacity, revealing a stale custom-metric scrape mismatch.
A Delayed Log Forwarder Buffer Hides a Checkout Service Error Spike During Incident Command
A fictional Daily Triage exercise: two Logging dashboards disagree during an incident, and the fix is verifying forwarder lag before changing severity.
Consumer Group Rebalancing Storm Stalls a Kafka Order Pipeline
A fictional Kafka consumer group suffers repeated rebalances. Broker and consumer evidence initially conflict — the reveal shows the real cause lies in JVM garbage-collection pauses.
Replication Lag Dashboards Mask a Stalled PostgreSQL Standby
A fictional PostgreSQL triage exercise: a replication dashboard reports healthy lag while a standby's WAL replay is actually stalled behind a long-running query.
A Branch-Based Cache Key Reintroduces a Patched CI/CD Dependency
A fictional triage exercise in which a CI/CD pipeline's cache key, built only from branch name and operating system, silently restores a pre-patch dependency tree on the main branch while feature branches appear correctly patched.
Split-Horizon DNS Answers Diverge After a Failed Zone Transfer
A fictional Networking & DNS triage exercise in which a secondary DNS server keeps serving a stale zone after a silent AXFR failure, producing conflicting answers for internal clients.
A Linux Log Service Fails to Write While Disk Space Appears Available
A fictional Daily Triage exercise in which a Linux log-forwarding service stops writing due to ENOSPC errors while df -h reports free disk space, revealing an inode-exhaustion fault caused by a logrotate misconfiguration.
Conflicting OpenTelemetry Alerts Mask a Silent Metrics Pipeline Drop
Two OpenTelemetry-driven alerts contradict each other during a fictional checkout-service incident. The reveal traces the conflict to a silently saturated metrics batch processor queue.
Cgroup v2 Page Cache Accounting Triggers Conflicting OOMKilled Alerts in a Kubernetes Checkout Service
A fictional Kubernetes triage exercise in which a checkout-service pod is OOMKilled while node-level dashboards report normal memory, and the reveal traces the contradiction to cgroup v2 page cache accounting.
HTTP 500 Spike Masks Log-Pipeline Backpressure
A fictional Incident Command triage exercise: an HTTP 500 spike looks like a deployment regression until log and disk evidence reveals log-pipeline backpressure as the real cause.
Kafka Consumer Group Falls Behind on One Partition
A fictional Kafka consumer group falls behind on one partition. Practise ruling hypotheses in or out with read-only evidence before recommending a safe fix.
CI/CD Pipeline Passes Yet Loses Its Build Artifact
A fictional CI/CD triage drill: a pipeline reports success while silently losing its build artifact, practising evidence-led diagnosis before any change.