Master Production
Troubleshooting
Every drill drops you into a system that is already broken: a live shell, an SLA clock running down, and grading on every command — including the ones that make it worse. Not a video of someone else fixing it.
Guided Paths
All 12 tracks →Role-shaped journeys that stack the modules in the order they actually build on each other.
DevOps Foundations
Shell, Linux internals, Git and containers — the ground floor everything else stands on.
- 1 Shell Scripting live
- 2 Linux Systems live
- 3 Git & Version Control live
- 4 Docker & Containers live
Kubernetes Operator
Scheduling, QoS evictions, DNS cascades and Helm packaging — run clusters that stay up.
- 1 Kubernetes live
- 2 Helm live
- 3 Ansible live
- 4 Jenkins & CI/CD live
Cloud & FinOps
Terraform state, AWS and GCP operations, then cutting the bill without cutting reliability.
- 1 Terraform live
- 2 AWS Operations live
- 3 GCP Operations live
- 4 Cloud Cost Optimization live
Individual Skills
231 lessons · 86h totalShell Scripting
Master bash scripting, command-line pipelines, input validation, positional arguments, looping, file parsing, and defensive programming for automation.
Linux Systems
Understand kernel diagnostics, process states, signals, memory hierarchy, system call tracing, networking diagnostics, and storage optimization.
Git & Version Control
Learn the object model first, then everything falls out of it: the three trees, branches as pointers, merge and rebase mechanics, remotes and refspecs, reflog recovery, bisect forensics, history rewriting and repository trust.
Docker & Containers
From namespaces and cgroups to a shipped image: layers and copy-on-write, Dockerfiles, multi-stage builds, volumes, the container network model, registries, Compose and hardening.
Kubernetes
From the reconciliation loop to production operations: workloads, scheduling, the pod network, RBAC, autoscaling, upgrades and etcd — pinned to Kubernetes 1.36.
Helm
Charts, values and releases: the object model behind helm upgrade, the template language in the forms you actually write, dependencies and OCI distribution, hooks and chart tests, and the upgrade failures that only appear on the second deploy.
Ansible
Master enterprise Ansible automation: Agentless SSH, Inventory design (group_vars/host_vars), Idempotent playbooks, Jinja2 templating, Variable precedence, Handlers, Task Blocks & Rescue error handling, Modular Roles & Galaxy, Ansible Vault secrets, and AWX / Ansible Tower job templates.
Terraform
From the declarative model to a production three-tier build: providers and resources, plan reading, the type system, state and remote backends, modules, workspaces, provisioners, CI/CD, policy as code, multi-cloud and a 100-error troubleshooting catalogue.
Jenkins & CI/CD
Delivery end to end: declarative Jenkins pipelines, shared libraries and ephemeral agents, then Google Cloud Build and Cloud Deploy, deployment strategies, supply-chain signing and a pipeline you can defend commit by commit.
AWS Operations
Run AWS the way it fails: the blast-radius model, IAM evaluation order, VPC routing and private connectivity, EC2 and EBS ceilings, RDS failover, EKS identity and capacity, then a landing zone you break on purpose.
GCP Operations
Master production GCP operations: Resource Hierarchy & Labels, IAM & Service Account delegation, Shared VPCs & Firewall Tags, GCE MIGs & SUD/CUD cost optimization, GCS WORM Lifecycles, GKE Pod IP & Workload Identity, Cloud Logging Sinks, and Pub/Sub & BigQuery analytics.
Cloud Cost Optimization
Cloud spend as an engineering discipline: the layer model, tagging and baselining, right-sizing without causing outages, CPU architecture migration, Spot and commitments, Karpenter and cluster cost tooling, storage, network and the observability bill.
Interactive Syllabus
Full dependency map →Featured War Rooms
All 12 rooms →Kubernetes Memory Eviction Cascade & CoreDNS Outage
midDiagnose a critical payment gateway latency spike caused by a node-level memory eviction chain reaction and silent DNS packet drops.
Production SSH/Bastion Outage Debugging
juniorDiagnose an intermittent SSH latency and connection timeout issue affecting engineers jumping from a bastion host to a production database VM.
Kubernetes War Room: Checkout API Pod Stuck in Pending
midA checkout API pod sits in Pending and payments stall. Work the scheduler evidence to the real constraint before reaching for more nodes.
Incident Replays
All 9 replays →A war room asks whether you can fix it. A replay walks the same outage decision by decision — including the plausible wrong turn at each fork, and what it costs you. Read these before an interview, when you need the reasoning and not just the fix.
Four Faults, One Host
The host serves production workloads; compounding CPU, DNS, and FD exhaustion drags every process on the box toward failure.
The Eviction Cascade
18% checkout failure rate during the peak window and a 12-minute transaction-processing degradation, amplified by a client retry storm.
The Runaway Bill
Budget breached and compounding month over month, with no attribution model to assign spend to the teams creating it — and every proposed fix carries the risk of degrading a production service to save money.
The Inherited Permission
Data pipeline jobs fail across the platform after a change that every check reported successful, and the failure looks like an application bug rather than an infrastructure one.
Featured Field Manuals
All 24 manuals →Karpenter & CastAI: Node Pool Design and Tool Selection
Designing Karpenter node pools for mixed real-time and batch workloads, and deciding when CastAI earns its place alongside them.
AWS Cost Optimization: A Healthcare Platform Case Study
A layered AWS cost optimization walkthrough on a healthcare platform: tagging, EC2 right-sizing, CPU architecture migration, storage and Savings Plans.
Masterclass: Designing Scalable, Low-Latency, Production-Grade AI Agent Systems
Designing agent systems that hold up under production load: DAG orchestration, a three-tier memory hierarchy, intent routing, speculative execution and TTFT optimisation.
Command Cheat Sheets
All 6 sheets →For when you are already in the incident and need the command, not the explanation. Ordered the way you actually work a problem, with the metrics that mislead called out.
AWS CLI Triage Cheat Sheet
The AWS CLI commands you actually run during an incident, in the order you run them — from 'which account am I in' to a quota that stopped a scale-out.
Bash Scripting Cheat Sheet
Quoting, expansion, strict mode, traps and the constructs that silently corrupt data — the reference for writing shell that survives production.
Git Recovery Cheat Sheet
Undoing the commit, the push, the rebase and the reset — including how to recover work that looks permanently gone.
kubectl Triage Cheat Sheet
The commands you actually run during a Kubernetes incident, in the order you run them — from a Pending pod to a dead node to an empty Service.
Linux Network Diagnosis Cheat Sheet
Walk a network fault layer by layer — reachability, routing, ports, DNS, TLS and packet capture — and prove whether latency lives in the network or the application.
Linux Performance Triage Cheat Sheet
A repeatable 60-second triage for a slow or unresponsive host — CPU, memory, disk and process state — plus the metrics that routinely mislead.
The shell is already running.
Every war room boots with the incident already in progress. Type real commands, get real output, and watch the SLA clock burn while you narrow it down.