Debug · Diagnose · Ship Under Pressure

Master Production Troubleshooting

Every drill drops you into a system that is already broken: a live shell, an SLA clock running down, and grading on every command — including the ones that make it worse. Not a video of someone else fixing it.

83 graded steps
231 hands-on lessons
0 videos to sit through
war-room — linux-ssh-port-outage
live outage
hover to pause
12 Timed War Rooms
24 Field Manuals
260 Interview Cards
9 Incident Replays

Guided Paths

All 12 tracks →

Role-shaped journeys that stack the modules in the order they actually build on each other.

Individual Skills

231 lessons · 86h total
🐚 available

Shell Scripting

Master bash scripting, command-line pipelines, input validation, positional arguments, looping, file parsing, and defensive programming for automation.

Bash SyntaxQuoting & ExpansionConditionals +10
26 lessons 583m
🐧 available

Linux Systems

Understand kernel diagnostics, process states, signals, memory hierarchy, system call tracing, networking diagnostics, and storage optimization.

ProcessesMemory & RAMSyscalls +7
15 lessons 265m
🌳 available

Git & Version Control

Learn the object model first, then everything falls out of it: the three trees, branches as pointers, merge and rebase mechanics, remotes and refspecs, reflog recovery, bisect forensics, history rewriting and repository trust.

Object modelThe three treesIndex & staging +18
19 lessons 393m
🐳 available

Docker & Containers

From namespaces and cgroups to a shipped image: layers and copy-on-write, Dockerfiles, multi-stage builds, volumes, the container network model, registries, Compose and hardening.

Containers vs VMsEngine internalsNamespaces & cgroups +17
14 lessons 313m
☸️ available

Kubernetes

From the reconciliation loop to production operations: workloads, scheduling, the pod network, RBAC, autoscaling, upgrades and etcd — pinned to Kubernetes 1.36.

Control planeWorkloadsScheduling +9
32 lessons 815m
⛵ available

Helm

Charts, values and releases: the object model behind helm upgrade, the template language in the forms you actually write, dependencies and OCI distribution, hooks and chart tests, and the upgrade failures that only appear on the second deploy.

Release modelChart anatomyRevisions & rollback +17
15 lessons 315m
🎭 available

Ansible

Master enterprise Ansible automation: Agentless SSH, Inventory design (group_vars/host_vars), Idempotent playbooks, Jinja2 templating, Variable precedence, Handlers, Task Blocks & Rescue error handling, Modular Roles & Galaxy, Ansible Vault secrets, and AWX / Ansible Tower job templates.

Agentless SSHansible.cfgStatic & Dynamic Inventories +15
10 lessons 240m
🏗️ available

Terraform

From the declarative model to a production three-tier build: providers and resources, plan reading, the type system, state and remote backends, modules, workspaces, provisioners, CI/CD, policy as code, multi-cloud and a 100-error troubleshooting catalogue.

IaC fundamentalsProviders & resourcesPlan reading +20
22 lessons 454m
⚙️ available

Jenkins & CI/CD

Delivery end to end: declarative Jenkins pipelines, shared libraries and ephemeral agents, then Google Cloud Build and Cloud Deploy, deployment strategies, supply-chain signing and a pipeline you can defend commit by commit.

CI vs CDDORA metricsBuild once, promote +23
18 lessons 397m
☁️ available

AWS Operations

Run AWS the way it fails: the blast-radius model, IAM evaluation order, VPC routing and private connectivity, EC2 and EBS ceilings, RDS failover, EKS identity and capacity, then a landing zone you break on purpose.

Blast radiusIAM evaluationSTS & IMDSv2 +17
20 lessons 445m
🚀 available

GCP Operations

Master production GCP operations: Resource Hierarchy & Labels, IAM & Service Account delegation, Shared VPCs & Firewall Tags, GCE MIGs & SUD/CUD cost optimization, GCS WORM Lifecycles, GKE Pod IP & Workload Identity, Cloud Logging Sinks, and Pub/Sub & BigQuery analytics.

Resource HierarchyOrg Policy guardrailsCloud IAM +28
28 lessons 649m
💰 available

Cloud Cost Optimization

Cloud spend as an engineering discipline: the layer model, tagging and baselining, right-sizing without causing outages, CPU architecture migration, Spot and commitments, Karpenter and cluster cost tooling, storage, network and the observability bill.

Layer modelTagging & attributionInventory & baselining +16
12 lessons 270m

Interactive Syllabus

Full dependency map →
SRE Learning Syllabus Syllabus Diagram
FoundationsKubernetesCloud ScalingAI OrchestrationLinux System Diagnos...Guide ManualjuniorSSH Port Outage TriageTimed ScenariojuniorSSH OSI Layer Deep-D...Timed ScenariomidNodePool Resource De...Guide ManualmidMemory Eviction & DN...Timed ScenariomidPending Pod Outage r...Timed ScenarioseniorEKS Karpenter ScalingGuide ManualmidAWS Cloud Cost Optim...Guide ManualmidAI Agent Systems Des...Guide ManualseniorProduction AI SystemsGuide Manualsenior
Click any node to study reference, drag to pan canvas, scroll to zoom

Featured War Rooms

All 12 rooms →

Incident Replays

All 9 replays →

A war room asks whether you can fix it. A replay walks the same outage decision by decision — including the plausible wrong turn at each fork, and what it costs you. Read these before an interview, when you need the reasoning and not just the fix.

Featured Field Manuals

All 24 manuals →

Command Cheat Sheets

All 6 sheets →

For when you are already in the incident and need the command, not the explanation. Ordered the way you actually work a problem, with the metrics that mislead called out.

No setup, no cluster, no cloud bill

The shell is already running.

Every war room boots with the incident already in progress. Type real commands, get real output, and watch the SLA clock burn while you narrow it down.

sre@warroom:~$ start linux-ssh-port-outage
provisioning sandbox…
✓ host prod-vm-ip up (10.0.1.18)
✓ incident timeline loaded — 4 phases
! SLA clock started: 15:00
sre@warroom:~$