Live incident drills

Incident Response War Rooms

The pager just went off. You get a shell, a countdown, and an outage that only ends when you find the actual root cause.

12 war rooms
6h of drills
4 senior-level
Difficulty junior · 1 mid · 7 senior · 4
junior SLA 20m

Production SSH/Bastion Outage Debugging

Diagnose an intermittent SSH latency and connection timeout issue affecting engineers jumping from a bastion host to a production database VM.

#linux#ssh
Enter →
mid SLA 30m

Kubernetes Memory Eviction Cascade & CoreDNS Outage

Diagnose a critical payment gateway latency spike caused by a node-level memory eviction chain reaction and silent DNS packet drops.

#kubernetes#incident-response
Enter →
mid SLA 30m

Production Outages Masterclass — Structured Debugging for DevOps/SRE

The five categories production outages fall into, and a structured debugging framework that replaces guessing with an ordered elimination of layers.

#kubernetes#aws
Enter →
mid SLA 30m

Kubernetes War Room: Checkout API Pod Stuck in Pending

A checkout API pod sits in Pending and payments stall. Work the scheduler evidence to the real constraint before reaching for more nodes.

#kubernetes#cloud-k8s
Enter →
mid SLA 30m

Kubernetes Outage Follow-Up: Structured Troubleshooting and RCA Writing

After the outage: a structured troubleshooting framework and how to write an RCA that survives review.

#kubernetes#aws
Enter →
mid SLA 30m

Production Outage: OSI-Layer Troubleshooting of a Bastion SSH Freeze

A bastion-to-production SSH freeze debugged layer by layer, from ICMP reachability up to the sshd configuration that caused it.

#kubernetes#aws
Enter →
mid SLA 30m

War Room Drill: Debugging a Production SSH/Bastion Outage

A production SSH/bastion outage worked end to end: symptom framing, blast radius, the login pipeline, and why the obvious fixes failed.

#kubernetes#cloud-k8s
Enter →
mid SLA 30m

War Room Drill: SSH Bastion-to-Prod Outage, an OSI Walkthrough

A bastion-to-production SSH freeze debugged layer by layer, from ICMP reachability up to the sshd settings that caused the stall.

#kubernetes#aws
Enter →
senior SLA 30m

War-Room Production Outage Drill: Kubernetes Memory Eviction Cascade & CoreDNS Outage

Pods evicting in a cascade while CoreDNS degrades underneath — working the QoS classes and node pressure signals back to the real trigger.

#kubernetes#cloud-k8s
Enter →
senior SLA 30m

Production Outage War Room: Triage Fundamentals

The five categories production outages fall into, and a triage sequence that narrows a vague “it is slow” to a subsystem fast.

#kubernetes#aws
Enter →
senior SLA 30m

War Room Drill: Kubernetes Pod Stuck in Pending

A checkout API pod stuck in Pending on EKS: ruling out the usual causes, finding three compounding root causes, and the scheduling concepts behind each.

#kubernetes#cloud-k8s
Enter →
senior SLA 30m

War Room Drill Follow-Up: Solution, Debugging Framework and RCA Writing

Resolving a Kubernetes scheduling outage, then writing the RCA — plus the cluster security scanning that should have caught it earlier.

#kubernetes#cloud-k8s
Enter →