The Green Illusion
Deployments Wedge Cluster-Wide While Every Dashboard Reports Perfect Health
The situation you’re stepping into
A managed Kubernetes cluster shared by several product teams. Deployments flow through CI, which applies manifests and waits for rollout to complete. The cluster runs the usual platform add-ons, including a handful of mutating and validating admission webhooks — policy enforcement, sidecar injection, image validation.
Admission webhooks are the part of this stack most teams install once and never think about again. They sit inside the write path of every single API request, which is precisely what makes this failure mode possible.
What the team observed
- New pods stick in
Pending. Deleted pods stick inTerminating. Both, simultaneously, across every namespace. - CI pipelines for multiple unrelated teams time out waiting for rollouts.
kubectl get nodes— allReady.- Cluster dashboards — CPU, memory, disk, network all normal. No alerts firing.
- The control plane reports healthy in the provider console.
kubectl get podsresponds fine. So doeskubectl describe.
Note the asymmetry, because it is the whole clue: reads are fast, writes are wedged. Every diagnostic command an engineer instinctively reaches for is a read, which is why the cluster feels responsive while nothing can actually change.
?
Decision Point 1 Pods stuck in Pending AND Terminating at the same time, cluster-wide, across unrelated namespaces. What does that combination rule out immediately, and what does it point at?
Pending is a scheduling concern. Terminating is a deletion concern. What single component is required for both, and what would have to be true for both to stall at once?
Commit to your answer, then reveal the responder’s move
→
Pods stuck in Pending AND Terminating at the same time, cluster-wide, across unrelated namespaces. What does that combination rule out immediately, and what does it point at?
Pending is a scheduling concern. Terminating is a deletion concern. What single component is required for both, and what would have to be true for both to stall at once?
Commit to your answer, then reveal the responder’s move →It rules out anything workload-specific and anything node-specific.
Reason about what each state requires:
Pendingmeans the pod object exists but no node has been assigned, or the assignment has not been acted on. That needs the scheduler to make a decision and write it back.Terminatingmeans a deletion timestamp is set but the object has not been removed. That needs controllers and kubelets to complete cleanup and write the finalisation.
A single namespace stuck is a workload problem. One node stuck is a node problem. Every namespace stuck in both directions at once has only one plausible common cause: the shared write path that all of it flows through — the API server, and everything the API server calls synchronously before it persists an object.
That reframing is the important move. The question stops being “what is wrong with these pods” and becomes “what is wrong with writes.”
Confirm it directly by timing a trivial write against a trivial read:
time kubectl get ns # read — fast
time kubectl label ns default probe=1 # write — hangs or takes seconds
?
Decision Point 2 Writes are slow, reads are fine, and the control plane is managed so you cannot read the API server process directly. Where do you look to find what a write is waiting on?
Every request the API server handles can be recorded with its duration and the stages it passed through.
Commit to your answer, then reveal the responder’s move
→
Writes are slow, reads are fine, and the control plane is managed so you cannot read the API server process directly. Where do you look to find what a write is waiting on?
Every request the API server handles can be recorded with its duration and the stages it passed through.
Commit to your answer, then reveal the responder’s move →The API server audit log. It records every request with its verb, resource, response code and — critically — duration and stage.
# Managed clusters expose these through the provider's logging backend.
# Filter for slow mutating requests:
# protoPayload.methodName =~ "create|update|patch|delete"
# AND duration > 1s
What you are looking for is the shape of the slowness, not a single slow request. Two patterns are worth distinguishing:
| Pattern | Points at |
|---|---|
| Writes slow across all resource types, reads fast | Something in the shared admission or persistence path |
| Writes slow for one resource type only | A controller or webhook scoped to that resource |
| Reads and writes slow | Storage layer — investigate etcd latency first |
If reads were also slow you would be looking at etcd, and the etcd request-duration metrics would be the next stop. They are not, so persistence is fine and the delay is happening before the write reaches storage.
The other signal available even without audit logs is the API server’s own request-duration metric, broken down by verb — the same asymmetry shows up there.
?
Decision Point 3 Audit logs show mutating requests consistently taking just over five seconds, then succeeding or failing. Why is that number itself the diagnosis?
A duration that clusters tightly on a round number is not congestion. Congestion is noisy. This is something expiring.
Commit to your answer, then reveal the responder’s move
→
Audit logs show mutating requests consistently taking just over five seconds, then succeeding or failing. Why is that number itself the diagnosis?
A duration that clusters tightly on a round number is not congestion. Congestion is noisy. This is something expiring.
Commit to your answer, then reveal the responder’s move →A tight cluster around a round number is a timeout, not load.
Genuine resource contention produces a spread — some requests fast, some slow, a long tail. When a large population of requests all take almost exactly the same duration and that duration is a round number, you are looking at a configured limit being reached and enforced.
Five seconds is the default admission webhook timeout. Every mutating request is calling out to a webhook, waiting the full timeout, and only then proceeding. The webhook is unreachable or unresponsive — the endpoint is gone, the backing pods are unhealthy, the Service has no endpoints, or a network policy is dropping the call.
Enumerate the webhooks and check whether their backends actually exist:
kubectl get validatingwebhookconfigurations
kubectl get mutatingwebhookconfigurations
# For each webhook's service reference — does it have endpoints?
kubectl -n <ns> get endpoints <webhook-service>
# No endpoints = every admission call waits the full timeout and then applies failurePolicy
?
Decision Point 4 Why did a five-second delay per request escalate into a cluster-wide freeze rather than merely making things slow? And why did failurePolicy make it worse?
The API server has a finite number of concurrent in-flight request slots, and controllers retry.
Commit to your answer, then reveal the responder’s move
→
Why did a five-second delay per request escalate into a cluster-wide freeze rather than merely making things slow? And why did failurePolicy make it worse?
The API server has a finite number of concurrent in-flight request slots, and controllers retry.
Commit to your answer, then reveal the responder’s move →Because the delay consumed a bounded resource, and the system’s own retries amplified it.
Three mechanisms compound:
1. In-flight request slots are finite. The API server caps concurrent mutating requests. Each one now occupies its slot for five seconds instead of milliseconds — a throughput collapse of roughly three orders of magnitude. Once every slot is held by a request waiting on the webhook, new writes queue behind them regardless of origin.
2. Controllers retry, which adds load. The scheduler, controller-manager, and every operator in the cluster react to a failed or slow write by retrying. Each retry takes another slot for another five seconds. The system’s normal self-healing behaviour becomes the amplifier — this is the deadlock: the busier it gets, the more it retries, and the more it retries the busier it gets.
3. failurePolicy: Fail turns a timeout into a rejection. With Fail, a webhook that does not answer causes the request to be rejected, which is the safe choice for a security policy and the dangerous one for availability. With Ignore, the request proceeds after the timeout — still slow, but not blocked. The policy choice determines whether an unavailable webhook degrades the cluster or stops it.
Why dashboards stayed green throughout: nothing they measured was wrong. Nodes were healthy, resource usage was normal, the API server process was alive and answering reads promptly. Admission latency was not instrumented at all — the failing component was the one nobody had a graph for.
Root cause
An admission webhook became unreachable, and every mutating API request paid its full timeout. With the default five-second timeout and a Fail failure policy, each write blocked for five seconds and was then rejected. Because admission runs inside the shared write path, this applied to every resource in every namespace — not just the workload the webhook cared about. Finite in-flight request slots filled with waiting calls, controller retries added further load, and the cluster reached a state where no write could complete: pods could neither be scheduled nor finalised, so they accumulated in Pending and Terminating simultaneously.
Monitoring reported healthy because every metric being collected was genuinely fine. The failure lived in a code path — synchronous admission — that had no instrumentation.
Resolution and prevention
# Immediate — stop the bleeding by removing the blocking call
kubectl patch validatingwebhookconfiguration <name> \
--type='json' -p='[{"op":"replace","path":"/webhooks/0/failurePolicy","value":"Ignore"}]'
# Or, if the webhook is not required for safety right now, remove it and reinstate later
kubectl delete validatingwebhookconfiguration <name>
# Then verify the write path recovers
time kubectl label ns default probe=2 # should return in milliseconds
kubectl get pods -A | grep -E 'Pending|Terminating' # backlog should drain
Prevention:
- Instrument admission latency as a first-class metric, with an alert threshold well below the timeout. This is the single change that would have turned a forty-five-minute outage into a page five minutes in.
- Set timeouts far below the default. One second is usually generous for a webhook that should answer in milliseconds. A shorter timeout bounds the blast radius of an unresponsive webhook.
- Choose
failurePolicydeliberately, per webhook.Failfor genuine security controls where bypass is unacceptable;Ignorefor convenience webhooks like sidecar injection or defaulting. The default is not a decision. - Scope webhooks narrowly with
namespaceSelector,objectSelectorand a minimalruleslist. A webhook that only needs to see one namespace should never be in the write path forkube-system. - Exclude system namespaces so a broken webhook cannot prevent the platform from recovering itself.
- Run webhook backends with multiple replicas, a PDB, and anti-affinity. A single-replica webhook backend with
failurePolicy: Failis a cluster-wide single point of failure wearing a security badge.
The generalisable rule: anything synchronous in the write path is a dependency of the entire cluster. Admission webhooks, and to a lesser extent finalizers, are the two places where installing a small component quietly makes every API operation depend on it.
Telling this story to a recruiter
Situation. A shared production Kubernetes cluster entered a partial freeze: pods hung in Pending and Terminating across every namespace at once, and release pipelines for multiple product teams stalled simultaneously. Every dashboard reported healthy — nodes Ready, resources normal, control plane green — so there was no failing signal to follow.
Task. Restore cluster write operations and unblock the deploy pipelines, with no obvious error to start from and multiple teams blocked.
Action. Recognised that Pending and Terminating stalling together across unrelated namespaces implicated the shared write path rather than any workload, and confirmed it by timing a trivial write against a trivial read — reads fast, writes wedged. Analysed API server audit logs and found mutating requests clustering tightly at just over five seconds, which identified a timeout being hit rather than resource contention. Traced that to an admission webhook whose backing service had no endpoints; with the default five-second timeout and a Fail failure policy, every write in the cluster was blocking and then being rejected, while controller retries amplified the queue into a deadlock. Patched the failure policy and timeout to restore the write path, then verified etcd latency was normal to confirm persistence had never been the problem, and cross-checked pipeline logs to prove deployments had been blocked by cluster state rather than application configuration.
Result. Cluster write operations and all team deploy pipelines restored within forty-five minutes. Instrumented admission webhook latency as a dedicated metric with alerting, closing the monitoring blind spot that had allowed a total write outage to present as a perfectly healthy cluster. Introduced an SLO for the admission path and tightened webhook timeouts, scoping and failure policies across the platform. No similar incident recurred in the following eighteen months.
What this demonstrates. Diagnosing from the absence of signal rather than the presence of an error; reading a latency distribution well enough to distinguish a timeout from congestion; and understanding the control plane deeply enough to know which components sit inside the synchronous write path.
Interview deep-dive: the full case study
Why “green dashboards” is the interesting part. Monitoring was not broken. Every metric it collected was accurate and every one of them was fine. The failure lived in a code path that had no metric attached — synchronous admission — so the monitoring system correctly reported that everything it could see was healthy. This is a category of outage that alerting cannot catch by tuning thresholds; it is only fixed by instrumenting the missing path. The lesson generalises well beyond webhooks: after any incident, the question worth asking is not “why did the alert not fire” but “was this failure mode observable at all”.
Why the read/write asymmetry is such a strong signal. Reads bypass admission entirely — they do not mutate state, so no webhook is consulted. Writes traverse authentication, authorisation, mutating admission, validating admission, then persistence. When reads are healthy and writes are not, the fault is almost certainly in a stage that only writes touch. Timing one of each is a five-second test that eliminates most of the search space.
Why the timing distribution matters more than the magnitude. Slowness caused by contention is noisy — a spread of durations with a long tail, because requests compete for a resource that frees unpredictably. Slowness caused by a timeout is tight, clustering just above a configured value, because every affected request waits the same fixed period. Learning to read that shape turns “the cluster is slow” into “something is timing out at five seconds” without any further investigation.
Why retries turned slowness into deadlock. This is the part that makes it an outage rather than a degradation. Kubernetes controllers are designed to retry — that is what makes them resilient. When the failure is capacity in the request path, retries consume the very resource that is scarce. The system’s self-healing mechanism becomes a positive feedback loop, and throughput collapses rather than degrading gracefully. Any system with bounded concurrency and automatic retry has this failure mode available to it.
The design lesson about failurePolicy. It encodes a genuine trade-off: Fail means an unavailable policy engine stops changes, which is correct for a control that must never be bypassed; Ignore means an unavailable webhook is skipped, which is correct for convenience features. Most teams never make the choice — they accept a default and inherit its consequences during an incident. The right practice is to decide per webhook, write down the reasoning, and make sure anything set to Fail is deployed with the redundancy that decision implies.