The Observability Blackout

Every Log Stream Goes Silent Mid-Incident, and the Pipeline Reports Healthy

SEV-1 Kubernetes ~40m

The situation you’re stepping into

A managed Kubernetes cluster running roughly forty services. Application containers write structured JSON to stdout. A log-forwarding DaemonSet on every node tails the container log files, and forwards to a central aggregation layer, which fans out to a hosted observability platform and a self-managed search cluster for long-term retention.

Nothing about this design is unusual. It is the shape most clusters end up with, and it has one property worth naming up front: it is a pipeline, and a pipeline can fail at any hop while every hop reports healthy.

You are already mid-investigation on an unrelated latency complaint when the logs stop arriving.

What the team observed

  • Dashboards that query logs go flat. Not error-flat — empty. The last data point is roughly twenty minutes old.
  • Every forwarder pod is Running, 1/1 Ready, zero restarts.
  • The aggregation layer is Running. Its own liveness endpoint returns 200.
  • The ingestion endpoint on the hosted platform reports no incident and no rate limiting.
  • Node CPU, memory and disk are unremarkable.
  • Applications are unaffected — they are writing to stdout exactly as before, and traffic is being served normally.
⚠️ Watch out

The reason this is a SEV-1 rather than an annoyance: it happened during an active investigation. The team lost the instrument they were using to measure the problem they were already working on, and every subsequent decision had to be made from metrics and guesswork.

?
Decision Point 1

Every component in the path reports healthy and no logs are arriving. Before touching any config, what is the single most useful thing to establish — and why is 'is the forwarder running?' the wrong first question?

A pipeline has an input, several hops, and an output. Running tells you a process exists. It does not tell you data is moving through it.

Commit to your answer, then reveal the responder’s move →
✓What the responder did & why

Establish where in the pipeline the data stops, by validating each hop independently rather than trusting the aggregate.

Running and Ready describe a process, not a flow. A forwarder whose output plugin cannot deliver will keep running, keep tailing files, keep accepting input, and keep reporting healthy — while buffering or discarding everything downstream of the break. That is the normal behaviour of a well-written forwarder, not a defect, and it is exactly why the health signal is useless here.

Walk the path in order and prove each link carries data:

# 1. Is the application actually emitting? (source of truth, bypasses the whole pipeline)
kubectl logs deploy/api --tail=20

# 2. Is the node-level forwarder reading those files?
kubectl -n logging logs ds/log-forwarder --tail=50

# 3. Is the aggregation layer receiving anything?
kubectl -n logging logs deploy/log-aggregator --tail=50

Step 1 matters more than it looks. If the application stopped logging, everything downstream is behaving correctly and you have a completely different problem.

?
Decision Point 2

Application stdout is fine. The forwarder's own log is mostly quiet but contains repeated retry and connection lines. What internal signal tells you definitively whether it is delivering, and what does a climbing buffer mean?

Forwarders expose their own metrics about records in versus records out. The gap between those two numbers is the whole story.

Commit to your answer, then reveal the responder’s move →
✓What the responder did & why

A forwarder exposes internal metrics about its own throughput — records read, records emitted, retries, dropped records, and current buffer usage. That endpoint is the instrument that tells you what Ready cannot.

kubectl -n logging port-forward ds/log-forwarder 2020:2020
curl -s localhost:2020/api/v1/metrics | jq

What you are looking for is the relationship between input and output:

Input risingOutput risingMeaning
yesyesPipeline healthy
yesflatDelivery is broken — this is your case
flatflatThe source stopped; look upstream at the app

Input climbing with output flat means the forwarder is reading fine and cannot hand off. Its buffer absorbs the difference — and a buffer is a fixed-size delay line, not a fix. Once it fills, the forwarder starts dropping, and depending on configuration it may drop silently.

That buffer is also why the failure appeared to start twenty minutes before you noticed: the pipeline kept delivering from buffer for as long as it could, so the visible outage lags the actual break.

?
Decision Point 3

The forwarder cannot deliver to the aggregator, yet both pods are Running and the aggregator's health endpoint returns 200. What class of failure produces exactly this, and why does a pod-level health check never catch it?

Think about what a long-lived TCP connection between two processes can do that neither process notices.

Commit to your answer, then reveal the responder’s move →
✓What the responder did & why

A broken forward connection — the TCP session between forwarder and aggregator is dead or half-open, while both processes remain alive and healthy by their own definition.

This is the specific failure class that defeats pod-level health checking:

  • The forwarder holds a socket it believes is open. Writes go into the kernel send buffer and are never acknowledged. From the process’s point of view it is working.
  • The aggregator is listening, healthy, and simply has one fewer client than it thinks. Its liveness endpoint has no opinion about connections that no longer exist.
  • A half-open connection — one side gone, the other unaware — survives until something forces the issue. Without keepalives, that can be a very long time.

Common triggers: an intermediate load balancer or network policy silently reaping idle connections, an aggregator restart the forwarder never noticed, or an MTU or firewall change on the path.

The diagnostic that settles it is to look at the socket itself rather than either process:

kubectl -n logging exec ds/log-forwarder -- ss -tnp | grep <aggregator-port>
# ESTAB with a large, non-draining Send-Q is the signature: queued and unacknowledged

A non-draining send queue on an ESTAB socket is the proof. The connection is established in name only.

?
Decision Point 4

You can restore delivery by restarting the forwarder DaemonSet. Why is that the wrong move here, and what do you do instead?

Restarting discards whatever is still in the buffer, and it destroys the evidence explaining why the connection broke.

Commit to your answer, then reveal the responder’s move →
✓What the responder did & why

Restarting works, and it costs you two things you cannot get back.

It discards the buffer. Everything the forwarder is still holding — including logs from the window you are actively investigating — is gone on restart. Those are precisely the records you most need.

It destroys the evidence. The half-open socket is the artifact that explains the failure. Restart and you have a working pipeline, no root cause, and a guaranteed recurrence.

The better sequence is to fix the connection handling in place:

  1. Reconfigure the output plugin with explicit retry limits and backoff so a failed delivery is retried rather than queued indefinitely.
  2. Enable TCP keepalives on the forward connection, so a dead peer is detected in seconds rather than never.
  3. Tune buffer sizing and the on-full behaviour deliberately — decide whether this pipeline should block or drop when full, rather than inheriting a default.
  4. Trigger a configuration reload rather than a pod restart, so the forwarder re-establishes its connection while keeping its buffered records.

The pipeline drains and the backlog delivers. No workload was restarted, and the buffered logs from the incident window survive.

Root cause

The forward connection between the node-level forwarder and the aggregation layer was silently broken, and nothing in the health model could see it. The forwarder continued reading container logs and queueing them into its buffer; the aggregator continued listening and reporting healthy; the socket between them was established but no longer carrying data. Records accumulated in buffer until it filled, then began dropping. Every liveness and readiness probe in the path passed throughout, because each one measured process health rather than data flow.

The failure was invisible for the same structural reason DNS saturation is invisible: the fault lived between the components, and health checks are defined on the components.

Resolution and prevention

# Immediate — restore delivery without discarding the buffer
kubectl -n logging edit configmap log-forwarder-config    # retry, backoff, keepalive, buffer policy
kubectl -n logging rollout restart ds/log-forwarder       # ONLY if a reload signal is unavailable

# Verify flow, not liveness
curl -s localhost:2020/api/v1/metrics | jq '.output'      # output count must be climbing

Prevention — make delivery itself a monitored signal:

  • Synthetic log injection. Emit a known marker line on a schedule, and assert it arrives in the destination within an expected window. This is the only check that tests the whole path end to end. If the marker stops arriving, the pipeline is broken regardless of what any component reports.
  • Alert on the input/output gap, not on pod health. A forwarder whose input rate exceeds its output rate for more than a minute is failing, whatever its probes say.
  • Alert on buffer utilisation with enough headroom to act before records are dropped.
  • Enable keepalives on every long-lived forward connection so half-open sockets are detected rather than waited on.
  • Treat log pipeline health as an SLI. It is infrastructure that other infrastructure depends on, and it deserves the same treatment as the services it observes.
🧭 Insight

The synthetic-injection check is the one that generalises. Any pipeline — logs, metrics, traces, events — can fail with every component healthy. The only reliable test is to put a known thing in one end and assert it comes out the other.


Telling this story to a recruiter

Situation. Centralised logging failed silently across an entire production cluster, during an active investigation into an unrelated issue. Every component in the pipeline — the per-node forwarders, the aggregation layer, the ingestion endpoint — reported healthy and Ready. Teams lost all log visibility at the worst possible moment and were forced to debug production blind.

Task. Restore log visibility end to end, without restarting application workloads and without discarding the buffered records covering the incident window.

Action. Validated each hop of the pipeline independently rather than trusting aggregate health — application stdout, node-level forwarder, aggregation layer, and ingestion endpoint. Used the forwarder’s internal throughput metrics to establish that input was climbing while output was flat, which localised the fault to delivery rather than collection. Inspected the socket directly and found an established connection with a non-draining send queue — a half-open TCP session that both processes still believed was healthy. Reconfigured the output plugin with retry, backoff, keepalives and an explicit buffer policy, and reloaded configuration rather than restarting pods, preserving the buffered records.

Result. Full log visibility restored across all services in about thirty-five minutes, with no workload restarts and no loss of the buffered incident-window logs. Introduced synthetic log injection with assertion monitoring so the pipeline is now validated end to end continuously, making log-delivery health a first-class SLI. Blind-debugging during incidents was eliminated as a failure mode, and mean time to resolution on subsequent incidents improved materially because the primary diagnostic signal became reliably available.

What this demonstrates. Reasoning about a distributed pipeline where every component lies by omission; knowing the difference between process health and data flow; and choosing the slower fix that preserves evidence over the faster one that destroys it.


Interview deep-dive: the full case study

Why the health model failed. Liveness and readiness probes are defined per component. They answer “is this process alive and able to accept work?” They cannot answer “is work moving between these processes?” — because that property does not belong to either component, it belongs to the link. A pipeline of N healthy components can carry zero data, and every probe will pass. This is the same structural blind spot that makes a saturated DNS service invisible to pod health, and the same one that makes a broken service-to-service dependency look like an application bug.

Why buffering delays detection. A forwarder buffers to absorb transient downstream slowness — that is correct and desirable behaviour. The side effect is that the visible start of an outage lags the actual break by however long the buffer lasts. When you reconstruct the timeline afterwards, you must work backwards from buffer size and ingest rate to find the real break time, or you will correlate the failure with the wrong change.

Why half-open connections are so persistent. TCP has no inherent liveness check. If one peer disappears without sending a FIN or RST — killed abruptly, or cut off by an intermediary that drops state — the other side’s socket remains ESTABLISHED indefinitely. It will only discover the truth when it writes enough data to exhaust the send window and hit a retransmission timeout, which can take a long time and may never resolve if the write rate is low. Keepalives exist precisely to bound this, and they are off by default in most stacks.

Why the restart was tempting and wrong. Restarting is the fastest path to a working pipeline, which is exactly why it gets chosen under pressure. It also guarantees you lose the buffered records covering the incident, and it removes the socket state that explains the failure — so you get a green dashboard, no root cause, and a recurrence. Under incident pressure, the discipline is to ask what evidence a given action destroys before taking it.

The generalisable lesson. Monitor the outcome, not the components. Synthetic injection with an assertion at the far end is the only check that validates a pipeline as a whole, and it is cheap: one known record, one scheduled assertion, one alert. Any observability pipeline without one is trusting that N independent health checks imply end-to-end delivery, and that implication does not hold.