A log shipping agent cannot reach its destination for 40 minutes — what are its options, and what does a disk-backed buffer change over an in-memory one?
answer
- The producer keeps going when the sink stops
- Memory, disk, stop reading, or drop
- Disk survives a restart, shares the node
- Refusing to read moves the risk to rotation
- Acknowledge late, and duplicates follow
basics
~20 sIt can buffer in memory, buffer to disk, stop reading and let pressure build, or drop. Disk survives an agent restart and holds far more, at the cost of node disk and I/O. Retry after a failed acknowledgement means duplicates.
solid answer
~40 sFour options, and every pipeline picks one whether or not anyone chose. **Buffer in memory**: fast, bounded by RAM, lost entirely if the agent restarts. **Buffer to disk**: survives restarts and holds far more, but consumes a node resource other workloads share. **Stop reading**: for file sources this means not advancing the read offset, so the risk moves to rotation deleting unread files; where the application writes into the agent, backpressure reaches the service and a synchronous logger can stall request threads. **Drop**: legitimate if you pick oldest or newest deliberately and count what went. Because the read offset only advances after the destination acknowledges, delivery is **at least once**: a batch acknowledged after the agent gave up on it is delivered twice. For the reader, log-line counts are therefore not event counts.
code
pseudocode · 9 linesbatch = reader.next(offset)
if destination.send(batch) == OK:
offset = batch.endOffset # only now is it delivered
else:
if not buffer.append(batch): # buffer is at its cap
choose_one_at_design_time:
stop_reading() # rotation may delete unread data
drop_oldest(); dropped++ # keeps the newest, loses the start
drop_newest(); dropped++ # keeps history, loses the incidentgo deeper
Know that a collection agent has to hold records somewhere when the destination is down, that its capacity is finite, and that what happens when it fills is a choice someone made.
Explain the difference between an in-memory and a disk-backed buffer — capacity, restart survival, and the node resource the disk one consumes — and why acknowledgements mean duplicates.
Show you have run one through an outage: rotation eating unread data while a reader is stalled, a logging call stalling request threads, and a fleet draining backlogs into a service that just recovered.
Own the guarantee as a platform contract: what the estate promises about loss and duplication, what readers may not assume, and how much node resource logging may take from the workloads.
When the destination goes away, the log pipeline does not stop producing. A parking-permit service still handles renewals at 1,240 requests a minute, still writes a line each, and the agent on each node has to do something with them for the next 40 minutes. What it does was decided by configuration nobody revisits until this moment. ## The four options 1. **Buffer in memory.** Records accumulate in the agent's process. This is fast and simple, is bounded by whatever memory limit the agent runs under, and evaporates completely if the agent is restarted or killed for exceeding that limit. On a busy node, memory buffers measure their capacity in minutes. 2. **Buffer to disk.** Records are written to files on the node and drained when the destination returns. Capacity goes from minutes to hours, and the buffer survives an agent restart or upgrade. 3. **Stop reading (backpressure).** The agent simply stops consuming, leaving the data where it is. Where the source is a file, this means not advancing the read offset. 4. **Drop.** Discard records once the buffer is full, choosing which end deliberately. ## Memory buffer versus disk buffer | | In-memory buffer | Disk-backed buffer | | --- | --- | --- | | Capacity | minutes of traffic, bounded by RAM | hours, bounded by a disk allowance | | Survives agent restart | no | yes | | Cost when idle | none | file space, plus write I/O on every batch | | Failure mode | agent killed for memory, everything lost | node disk fills, harming co-located workloads | | Effect on latency | none | small, from writing before sending | The disk buffer's failure mode is the one people miss: the buffer is on a resource shared with everything else on the node. An unbounded queue that fills the node's disk turns a destination outage into a node outage, so its allowance must be capped explicitly, and what happens at that cap is the real design decision. ## Backpressure has a destination too Backpressure is not a way of avoiding a decision; it moves the problem to whoever is upstream. - **With file sources**, refusing to advance the read offset leaves the lines in the container's log files. Those files rotate on size and old segments are deleted, so a long enough stall means the agent's next read finds that the data it was waiting on has been removed. This is the classic **silent** loss: nothing errored, records simply never existed downstream. - **When the application writes directly into the pipeline** — over a socket, or through a shipping library in-process — backpressure reaches the service. A synchronous logging call that blocks turns a log-destination outage into a latency incident, and then into an availability incident when the request threads are all parked in a logging call. Logging must degrade before the service does: a non-blocking appender with a bounded queue that drops, and a counter for what it dropped. ## At least once, and what it means to the reader The read offset — or the acknowledgement of a batch — only advances when the destination confirms receipt. If the destination processed a batch and the acknowledgement was lost, the agent retries and the records land twice. That is **at-least-once** delivery, and it is the honest default: the alternative, advancing before confirmation, silently loses data instead. For whoever reads the logs later, that has consequences: - **Counting log lines is not counting events.** Seeing three copies of an error may mean one failure delivered three times. - **Deduplication needs identity the emitter supplied.** Timestamp plus text is not identity, because two identical retries of the same real operation look the same as one duplicated record. If exact counts matter — evidence a regulator will ask for — the emitter has to write a unique event identifier. - **Order is not preserved across a retry.** A retried batch arrives after records produced later, so the store's arrival time diverges from the event time. Query on the emitted timestamp, and be careful with anything that assumes monotonic ingestion. ## Draining without causing the second outage When the destination comes back, an agent holding 40 minutes of records will try to ship them as fast as it can, and a fleet of agents doing that simultaneously is a thundering herd aimed at a service that has just recovered. Two habits prevent the follow-on incident: cap the drain rate so recovery traffic is a bounded multiple of normal ingest, and jitter retry backoff so the fleet does not reconnect in lockstep. ## What to have decided in advance 1. Buffer type and an explicit **size cap**, on a disk allowance that is not the same space the workloads need. 2. **What happens at the cap** — drop oldest, drop newest, or block — decided per source rather than left to a default. 3. A **dropped-records counter**, exported as a metric and alerted on, so a gap is visible without reading the logs you have just lost. 4. A **drain rate limit** and jittered backoff. 5. An agreement with readers that the pipeline is at-least-once, so nobody builds a count that assumes otherwise.
- Why can refusing to read be worse than dropping records outright?Because the loss becomes invisible. A drop is counted and can be alerted on; a stalled reader looks healthy while the container's log files rotate underneath it and old segments are deleted. You then have neither the records nor a number telling you how many are missing, and the gap is only discovered when someone searches for a line that should be there.
- The destination returns and immediately falls over again. What happened?Every agent in the fleet started draining its backlog at once, so the recovered service received a multiple of normal ingest from hundreds of senders simultaneously. Cap each agent's drain rate to a bounded multiple of its steady-state throughput and jitter the reconnect backoff, so recovery is spread over minutes instead of arriving as one synchronised burst.
- A regulator wants an exact count of a specific operation over six months. Can log records give it?Not on their own, under at-least-once delivery. Retries produce duplicates that are indistinguishable from genuine repeats unless the emitting process wrote a unique identifier per event, in which case the count is a distinct count over that field. Without one, treat the record stream as evidence that something happened, and take exact counts from a system with transactional identity.
saying these in an interview costs you the question
- Assumes an agent simply waits and loses nothing
- Sizes a disk buffer without capping its space
- Lets a logging call block request threads
- Thinks retries give exactly-once delivery
- Counts log lines as if each were one event
- Has no metric for records dropped or buffered