Incident Replay
Real production incidents replayed as a story you work through. Each drill sets the scene and the stakes, then walks the investigation as a sequence of decision points — commit to what you'd check first, then reveal what the responder actually did, and why.
Four Faults, One Host
A single production Linux host goes sluggish: load average through the roof, every hostname lookup crawls, and apps start failing with 'Too many open files'. Four independent faults were injected at the same time — a CPU storm, a DNS thread storm, a resolver misconfiguration, and a file-descriptor leak — and only a disciplined, layer-by-layer sweep untangles them.
The host serves production workloads; compounding CPU, DNS, and FD exhaustion drags every process on the box toward failure.
The Eviction Cascade
A payments namespace starts shedding pods at peak — checkout restarts, scheduled jobs die mid-run — with no deployment, no config change, and every node showing Ready. The culprit is a shared node pool with no isolation between a revenue-critical service and a hungry batch workload.
18% checkout failure rate during the peak window and a 12-minute transaction-processing degradation, amplified by a client retry storm.
The Silent DNS Saturation
A payments platform starts throwing intermittent 502s and 20x latency at peak — yet every pod is Running, no pod is OOMKilled or Evicted, every node is Ready, and nothing deployed. The failure lives in the one dependency that sits in front of every service call: cluster DNS.
~3,000 transactions/minute at peak; revenue-impacting SEV-1 as checkout latency jumps from ~90 ms to over 1.8 s.
The Bastion Mirage
Thirty engineers, one bastion, one private production host — and a failure that only some people can reproduce, only some of the time. A blast-radius and OSI-layer story where the 'obvious' single cause is a mirage hiding four independent faults.
Peak reporting day; payment-analytics backend at risk of a 1-hour SLA breach if engineers can't reach prod to remediate.
The Green Illusion
Pods hang in Pending and Terminating across every namespace, release pipelines stall for multiple teams at once, and the cluster's own monitoring insists nothing is wrong. The write path through the API server is queueing behind something nobody instrumented.
Every team's deploy pipeline is frozen simultaneously — no releases, no rollbacks, no hotfixes — while the cluster reports healthy and the on-call has no failing signal to follow.
The Inherited Permission
A migration from node-level cloud credentials to per-workload identity completes cleanly. Nodes are healthy, pods are Running, the new role is assumed successfully — and every application that touches object storage starts failing with access denied. The permission that vanished was never written down anywhere.
Data pipeline jobs fail across the platform after a change that every check reported successful, and the failure looks like an application bug rather than an infrastructure one.
The Observability Blackout
Centralised logging stops delivering across an entire cluster while every component in the path — DaemonSet pods, aggregator, ingestion endpoint — reports Running and Ready. Teams lose all log visibility in the middle of debugging a live production issue, and have to debug the thing they debug with.
Blind debugging during an active incident. Every team loses the primary signal they use to reason about production, and the outage clock keeps running on whatever they were already investigating.
The Rollout Freeze
A routine rolling update of a revenue-critical checkout service wedges: two replicas go healthy, the third sits Pending forever, and the rollout freezes. No crashes, no failing probes, no node problems — and node drains start hanging too. A pure structured-debugging exercise in a cluster that looks flawless.
Checkout runs at reduced replica count under production load, the deploy pipeline is frozen (no rollback, no hotfix), and node maintenance is blocked — a slow drift toward a larger outage.
The Runaway Bill
Monthly spend across two clouds breaches budget and keeps climbing while traffic is flat. Nothing is down, no alert fires, and no team recognises the resources driving the increase. A layered cost engagement where the hardest problem is attribution, and the biggest risk is causing an outage while trying to save money.
Budget breached and compounding month over month, with no attribution model to assign spend to the teams creating it — and every proposed fix carries the risk of degrading a production service to save money.