kubectl Triage Cheat Sheet

The commands you actually run during a Kubernetes incident, in the order you run them — from a Pending pod to a dead node to an empty Service.

Kubernetes command reference

The 60-second sweep

Run these four before forming any theory. Most incidents are visible here.

# What is actually broken, cluster-wide
kubectl get pods -A --field-selector=status.phase!=Running

# Recent events, newest last — the single highest-signal command
kubectl get events -A --sort-by=.lastTimestamp | tail -30

# Are the nodes healthy?
kubectl get nodes -o wide

# Is the control plane answering?
kubectl get --raw='/readyz?verbose'

Events are namespaced and expire after one hour by default. If an incident started earlier than that, the events are already gone — go to the controller logs instead.

Pod stuck in Pending

Pending means the scheduler has not placed it. The reason is always in the events.

kubectl describe pod POD -n NS | sed -n '/Events:/,$p'
MessageCause
Insufficient cpu / memoryNo node has enough allocatable capacity
node(s) had untolerated taintNeeds a toleration, or the node is cordoned
didn't match Pod's node affinityAffinity rules exclude every node
pod has unbound immediate PersistentVolumeClaimsNo PV, or the StorageClass cannot provision
node(s) exceed max volume countPer-node CSI attach limit reached
# What is actually free, as the scheduler sees it
kubectl describe node NODE | sed -n '/Allocated resources/,/Events/p'

# Requests are what the scheduler uses — not usage
kubectl get pods -A -o custom-columns=\
NS:.metadata.namespace,NAME:.metadata.name,CPU:.spec.containers[*].resources.requests.cpu

The scheduler places pods on requests, never on actual usage. A node at 5% CPU can still be unschedulable if its requests are fully committed.

CrashLoopBackOff

The container starts and exits repeatedly. Read the previous container’s logs — the current one is usually still starting.

kubectl logs POD -n NS --previous
kubectl logs POD -n NS -c CONTAINER --previous   # multi-container pod

# Exit code and reason
kubectl get pod POD -n NS -o jsonpath='{.status.containerStatuses[*].lastState.terminated}'
Exit codeMeaning
0Exited cleanly — usually a missing long-running process
1Application error, read the logs
137SIGKILL — almost always OOMKilled, or a failed liveness probe
139SIGSEGV — segfault
143SIGTERM — graceful shutdown that never completed

137 with reason: OOMKilled means the memory limit was hit. Raise the limit or fix the leak; restarting changes nothing.

Service returns nothing

A Service with no endpoints is a selector problem nine times out of ten.

kubectl get endpointslices -n NS -l kubernetes.io/service-name=SVC

# Compare these two — they must match exactly
kubectl get svc SVC -n NS -o jsonpath='{.spec.selector}'
kubectl get pods -n NS --show-labels

Empty endpoints means either no pod matches the selector, or matching pods are not Ready. Readiness controls Service membership — a failing readiness probe silently removes a pod from the load balancer while leaving it Running.

# DNS resolution from inside the cluster
kubectl run -it --rm dnstest --image=busybox:1.36 --restart=Never -- \
  nslookup SVC.NS.svc.cluster.local

Node problems

kubectl describe node NODE | sed -n '/Conditions:/,/Addresses:/p'
ConditionMeaning
MemoryPressurekubelet will start evicting pods
DiskPressureImage garbage collection, then eviction
PIDPressureToo many processes on the host
NotReadykubelet stopped reporting — check the kubelet itself
# Drain safely before maintenance
kubectl drain NODE --ignore-daemonsets --delete-emptydir-data

# Drain hanging? A PDB is blocking it
kubectl get pdb -A

A PDB with ALLOWED DISRUPTIONS: 0 will block a drain forever. That is the single most common cause of a stuck node upgrade.

Resource pressure

kubectl top nodes
kubectl top pods -A --sort-by=memory

# Requests vs allocatable — the ratio that predicts scheduling failure
kubectl describe nodes | grep -A5 "Allocated resources"

CPU limits cause throttling, not eviction, and throttling never appears as an OOMKill. If latency is bad but memory looks fine, check throttling:

kubectl exec POD -n NS -- cat /sys/fs/cgroup/cpu.stat | grep throttled

RBAC denials

# Can I do this?
kubectl auth can-i create deployments -n NS

# Can that ServiceAccount do this? — the one that matters during an incident
kubectl auth can-i list secrets -n NS \
  --as=system:serviceaccount:NS:SA_NAME

kubectl auth can-i --list --as=system:serviceaccount:NS:SA_NAME

Getting inside a container

kubectl exec -it POD -n NS -- sh

# Distroless image with no shell — attach a debug container instead
kubectl debug -it POD -n NS --image=busybox:1.36 --target=CONTAINER

# Debug a node without SSH
kubectl debug node/NODE -it --image=busybox:1.36

Rollouts

kubectl rollout status deploy/NAME -n NS --timeout=120s
kubectl rollout history deploy/NAME -n NS
kubectl rollout undo deploy/NAME -n NS            # previous revision
kubectl rollout undo deploy/NAME -n NS --to-revision=3

A rollout that hangs is usually the new pods failing readiness — check the new ReplicaSet’s pods, not the Deployment.

Evidence to capture before it expires

Events age out after an hour and deleted pods take their logs with them. Capture first, diagnose second.

NS=production
kubectl get events -n $NS --sort-by=.lastTimestamp > events.txt
kubectl describe pod POD -n $NS                   > describe.txt
kubectl logs POD -n $NS --previous --timestamps   > logs.txt
kubectl get pod POD -n $NS -o yaml                > pod.yaml