skip to content

During a burst, an hour of a container's captured output is missing centrally although the collector never went down — what happened?

level: seniorimportance: should knowfreq 48%

answer

  1. the node copy is bounded
  2. cap times files kept
  3. deleted under a lagging reader
  4. lost before ingest, not after
  5. alert on lag against retained window

basics

~20 s

The node's copy of the captured output is bounded by a size cap and a number of files kept. Under a burst the workload outran the collector, and the oldest files were deleted before it read them — those lines never entered the pipeline at all.

solid answer

~50 s

The runtime writes captured output to node-side files under a rotation cap: a maximum size per file and a maximum number of files kept, so a container's output on the node is bounded. The collector reads those files and ships onward. When emission outruns shipping, the collector's position falls behind the oldest retained file, and rotation deletes that file underneath it. Those lines are gone before ingest, so nothing central can replay them and no retention setting recovers them — they were never there. The tell is a contiguous gap for the noisiest workload during its busiest window, with the collector healthy throughout. The fix is to make the retained window on the node exceed the collector's worst-case lag, cut emission at the source, or both — and to alert on the lag approaching that window.

code

pseudocode · 16 lines
pseudocode
maxFileBytes = 10_000_000          # per capture file
filesKept    = 3
retainedBytes   = maxFileBytes * filesKept        # 30_000_000
emitBytesPerMin = 2_000_000
retainedMinutes = retainedBytes / emitBytesPerMin # 15

on captured_line(line):
    append(currentFile, line)
    if size(currentFile) >= maxFileBytes:
        rotate()                                   # currentFile becomes newest kept
        if count(keptFiles) > filesKept:
            delete(oldestKeptFile)                 # unread lines in it are gone

collectorLagMinutes = 20
if collectorLagMinutes > retainedMinutes:           # 20 > 15
    lostMinutes = collectorLagMinutes - retainedMinutes   # 5

go deeper

for a junior

Know that the captured copy on the node is bounded: a size cap and a limited number of files kept, with the oldest deleted. A container's log history on the node is minutes, not days.

for a middle

Compute the retained window — cap times files kept, divided by the emission rate — and explain why a collector that falls behind it loses lines permanently rather than catching up later.

for a senior

Diagnose it from the evidence: a contiguous gap for the noisiest workload, ending when the burst did, with a healthy collector throughout. Then name the lever you would pull and the lag alert you would add.

for a principal

The standard you set is the one nobody has: a stated retained window per workload class, measured against worst-case collector lag, with emission budgets that make it hold. Decide what evidence loss during a burst is acceptable, because some always is.

## Where the captured copy lives before it is shipped What the runtime reads from a container's output streams is written to files the node manages on the node's own disk. Those files are not storage; they are a **staging buffer with a hard ceiling**, because a node cannot let one talkative workload consume its disk. The ceiling has two parts: - a maximum size for the current file, after which it is rotated and a new one started; - a maximum number of rotated files kept, after which the oldest is deleted. Multiply them and you have the retained bytes. Divide by the workload's emission rate and you have the retained *time* — the real number, and the one nobody knows for their own services. ## The arithmetic of the gap Take a cap of 10 MB per file with 3 files kept: **30 MB retained**. A gateway under load emitting 2 MB per minute therefore has **15 minutes of history on the node** — 30 divided by 2. If the collector is 20 minutes behind, it is trying to read from a position that rotation deleted 5 minutes ago, and **5 minutes of lines are gone** before anything read them. Raise the burst to 6 MB per minute and the retained window collapses to 5 minutes, which is shorter than many restarts. The lag itself comes from the ordinary causes: shipping throughput below peak emission, the destination applying back-pressure, the collector restarting and re-reading, or simply many containers bursting on one node at once. ## Why nothing downstream can fix it This is the distinction that decides the whole diagnosis: **the loss happened before ingest.** | failure | where the lines are | what recovers them | |---|---|---| | destination unreachable for 40 minutes | still on the node, or in the collector's buffer | the collector retries and ships them when it can | | collector process restarts | still on the node, if within the retained window | it resumes from its recorded position | | node rotation deletes unread files | deleted from the node, never in the pipeline | nothing — no retry, no retention change, no replay | A disk-backed buffer on the collector protects the middle rows: it keeps lines the collector has already read when the destination will not take them. It does nothing for the bottom row, where the collector never read the lines in the first place. Extending retention in the central store is the same mistake one layer further on — it governs how long ingested records are kept, and these were never ingested. ## Reading the evidence The signature of rotation loss is specific, and it distinguishes it from the failures it gets confused with: 1. **The gap is contiguous**, not sparse — a stretch of wall-clock time with nothing from either stream, rather than lines with holes in them. 2. **It ends at the moment the burst subsided**, because that is when the collector caught up. 3. **It affects the noisiest workloads on that node**, and workloads with modest output on the same node are untouched, since each container's retained window is its own emission rate against the same cap. 4. **The collector was healthy the whole time** — which is exactly why the team looks past the real cause. Its own health says nothing about whether the files it wanted still existed. If instead everything on the node is missing across all workloads, or the node stopped accepting work, you are looking at the node running short of disk rather than at one workload's rotation cap. That is a different mechanism, with different symptoms and a different owner. ## What actually helps - **Size the retained window against worst-case lag, not against average emission.** The question is not "how much do we log" but "how long can the collector be behind before deletion overtakes it", and the answer must be measured during a burst. - **Measure the lag and alert on it approaching the window.** Collector-is-running is the wrong signal; collector-is-N-minutes-behind-and-the-node-holds-M is the right one. This is the single highest-value change, because the loss is otherwise invisible until someone goes looking for the lines. - **Cut emission at the source for the workloads that burst.** A debug-level flood turned on during an incident can shrink a fifteen-minute window to two and destroy the evidence for the incident it was turned on for — a genuinely nasty feedback loop. - **Make shipping throughput exceed peak emission**, not mean emission; the buffer is bounded in bytes, so the mean is not the constraint. Think of the node's files as a conveyor belt of fixed length: anything that reaches the far end before the reader picks it up falls off. Lengthening the belt buys time; it never puts back what already fell.

  • Would a larger disk-backed buffer on the collector have saved these lines?
    No. That buffer holds lines the collector has already read, so it covers a destination that will not accept them. Here the collector had not read them yet and rotation deleted the file underneath it. The lever is the retained window on the node, or less emission, or faster shipping.
  • How do you find out your real retained window without waiting for the next incident?
    Measure emission in bytes per minute for the workload at peak, and read the cap the node applies — retained bytes divided by peak rate is the window in minutes. Then compare it with the collector's observed lag during that same peak. If the two are close, you are already losing lines occasionally.
  • Why does turning up detail during an incident sometimes make the evidence worse?
    The retained window is bytes divided by emission rate, so multiplying output divides the window. A fifteen-minute buffer can become two, and the collector then falls behind it under the same load — so the extra detail you enabled is deleted before shipping, along with the ordinary lines you were relying on.

saying these in an interview costs you the question

  • Concludes the collector must have crashed because lines are missing
  • Expects a longer retention setting centrally to recover unshipped lines
  • Thinks a destination-side buffer protects against node rotation
  • Assumes the node keeps captured output until something reads it
  • Treats a contiguous gap and scattered missing lines as the same failure