kubectl Triage Cheat Sheet
The commands you actually run during a Kubernetes incident, in the order you run them — from a Pending pod to a dead node to an empty Service.
The 60-second sweep
Run these four before forming any theory. Most incidents are visible here.
# What is actually broken, cluster-wide
kubectl get pods -A --field-selector=status.phase!=Running
# Recent events, newest last — the single highest-signal command
kubectl get events -A --sort-by=.lastTimestamp | tail -30
# Are the nodes healthy?
kubectl get nodes -o wide
# Is the control plane answering?
kubectl get --raw='/readyz?verbose'
Events are namespaced and expire after one hour by default. If an incident started earlier than that, the events are already gone — go to the controller logs instead.
Pod stuck in Pending
Pending means the scheduler has not placed it. The reason is always in the events.
kubectl describe pod POD -n NS | sed -n '/Events:/,$p'
| Message | Cause |
|---|---|
Insufficient cpu / memory | No node has enough allocatable capacity |
node(s) had untolerated taint | Needs a toleration, or the node is cordoned |
didn't match Pod's node affinity | Affinity rules exclude every node |
pod has unbound immediate PersistentVolumeClaims | No PV, or the StorageClass cannot provision |
node(s) exceed max volume count | Per-node CSI attach limit reached |
# What is actually free, as the scheduler sees it
kubectl describe node NODE | sed -n '/Allocated resources/,/Events/p'
# Requests are what the scheduler uses — not usage
kubectl get pods -A -o custom-columns=\
NS:.metadata.namespace,NAME:.metadata.name,CPU:.spec.containers[*].resources.requests.cpu
The scheduler places pods on requests, never on actual usage. A node at 5% CPU can still be unschedulable if its requests are fully committed.
CrashLoopBackOff
The container starts and exits repeatedly. Read the previous container’s logs — the current one is usually still starting.
kubectl logs POD -n NS --previous
kubectl logs POD -n NS -c CONTAINER --previous # multi-container pod
# Exit code and reason
kubectl get pod POD -n NS -o jsonpath='{.status.containerStatuses[*].lastState.terminated}'
| Exit code | Meaning |
|---|---|
0 | Exited cleanly — usually a missing long-running process |
1 | Application error, read the logs |
137 | SIGKILL — almost always OOMKilled, or a failed liveness probe |
139 | SIGSEGV — segfault |
143 | SIGTERM — graceful shutdown that never completed |
137 with reason: OOMKilled means the memory limit was hit. Raise the limit or
fix the leak; restarting changes nothing.
Service returns nothing
A Service with no endpoints is a selector problem nine times out of ten.
kubectl get endpointslices -n NS -l kubernetes.io/service-name=SVC
# Compare these two — they must match exactly
kubectl get svc SVC -n NS -o jsonpath='{.spec.selector}'
kubectl get pods -n NS --show-labels
Empty endpoints means either no pod matches the selector, or matching pods are not Ready. Readiness controls Service membership — a failing readiness probe silently removes a pod from the load balancer while leaving it Running.
# DNS resolution from inside the cluster
kubectl run -it --rm dnstest --image=busybox:1.36 --restart=Never -- \
nslookup SVC.NS.svc.cluster.local
Node problems
kubectl describe node NODE | sed -n '/Conditions:/,/Addresses:/p'
| Condition | Meaning |
|---|---|
MemoryPressure | kubelet will start evicting pods |
DiskPressure | Image garbage collection, then eviction |
PIDPressure | Too many processes on the host |
NotReady | kubelet stopped reporting — check the kubelet itself |
# Drain safely before maintenance
kubectl drain NODE --ignore-daemonsets --delete-emptydir-data
# Drain hanging? A PDB is blocking it
kubectl get pdb -A
A PDB with ALLOWED DISRUPTIONS: 0 will block a drain forever. That is the
single most common cause of a stuck node upgrade.
Resource pressure
kubectl top nodes
kubectl top pods -A --sort-by=memory
# Requests vs allocatable — the ratio that predicts scheduling failure
kubectl describe nodes | grep -A5 "Allocated resources"
CPU limits cause throttling, not eviction, and throttling never appears as an OOMKill. If latency is bad but memory looks fine, check throttling:
kubectl exec POD -n NS -- cat /sys/fs/cgroup/cpu.stat | grep throttled
RBAC denials
# Can I do this?
kubectl auth can-i create deployments -n NS
# Can that ServiceAccount do this? — the one that matters during an incident
kubectl auth can-i list secrets -n NS \
--as=system:serviceaccount:NS:SA_NAME
kubectl auth can-i --list --as=system:serviceaccount:NS:SA_NAME
Getting inside a container
kubectl exec -it POD -n NS -- sh
# Distroless image with no shell — attach a debug container instead
kubectl debug -it POD -n NS --image=busybox:1.36 --target=CONTAINER
# Debug a node without SSH
kubectl debug node/NODE -it --image=busybox:1.36
Rollouts
kubectl rollout status deploy/NAME -n NS --timeout=120s
kubectl rollout history deploy/NAME -n NS
kubectl rollout undo deploy/NAME -n NS # previous revision
kubectl rollout undo deploy/NAME -n NS --to-revision=3
A rollout that hangs is usually the new pods failing readiness — check the new ReplicaSet’s pods, not the Deployment.
Evidence to capture before it expires
Events age out after an hour and deleted pods take their logs with them. Capture first, diagnose second.
NS=production
kubectl get events -n $NS --sort-by=.lastTimestamp > events.txt
kubectl describe pod POD -n $NS > describe.txt
kubectl logs POD -n $NS --previous --timestamps > logs.txt
kubectl get pod POD -n $NS -o yaml > pod.yaml