skip to content

Your only input is a push feed that keeps no history. How do you make a job over it recoverable, and what does that cost?

level: seniorimportance: should knowfreq 52%

answer

  1. a forgetful source caps recovery
  2. create the re-readable input yourself
  3. the boundary moves, not vanishes
  4. keep the receiving hop trivial
  5. acknowledge only after the append is durable

basics

~20 s

Put a durable, re-readable copy in front of the feed and have the job read that instead. Recoverability then starts at the copy, so the unrecoverable segment moves to the trivial hop that writes it — it shrinks, it does not vanish.

solid answer

~50 s

A source that hands each record over once and forgets it cannot support recovery, because recovery is reprocessing and there is nothing to reprocess from. The standard repair is to stop treating the feed as the job's input: a small component receives from the feed and appends every record, untouched, to durable storage that a reader can resume from at a named position, and the real job reads that copy. The consequence is that the boundary of recoverability moves rather than disappears — whatever the receiving hop holds when it dies is still gone. So you make that hop as close to trivial as possible: no parsing that can fail on content, no enrichment, no business logic, nothing that makes it slow or likely to crash. Where the feed supports being told a record was taken, the hop can record a write-ahead record of intent before acknowledging, which narrows the loss further but never to zero.

go deeper

for a junior

Know the shape of the answer: if the source cannot be re-read, write the records somewhere durable first and have the job read from there instead of straight from the feed.

for a middle

Explain why the job's recovery models all need re-readable input, and how the inserted copy supplies it by letting a reader resume at a named position after a restart.

for a senior

Say where the unrecoverable segment is once your copy exists, how big it is, and which properties of the receiving hop — no logic, fast restart, rare deploys — keep it small.

for a principal

Weigh the extra hop's latency, bill and operational surface against the value of being able to reprocess at all, and decide whether the sender should simply write durably in the first place.

## A source that forgets sets a hard ceiling Recovery is reprocessing, so a job's ability to survive a failure is bounded by what its input can give back. A push feed, a socket, a device stream or an in-memory hand-off gives one look at each record. Everything in flight when a machine dies is simply gone, and no engine mechanism changes that: **a recovery point** — a durable copy of everything the job would otherwise lose, written while it keeps running — preserves the job's own accumulated state, but it cannot preserve records that never reached durable storage anywhere. This is the ceiling to state plainly in an interview before proposing anything: *with this input, the job can be made available across failures but not correct across them.* ## The repair: a durable copy in front Make the job's real input **a retained, re-readable input** — a source that keeps its records for a while and lets a reader resume from a position it names — by creating one: 1. A receiving component subscribes to the feed and appends each record, byte-for-byte, to durable storage that supports resuming at a named position. 2. The processing job's source becomes that storage, not the feed. 3. The job records its **recorded read position** — the durable marker of how far it has got — against that storage, and a restart resumes there. The records are now re-readable, so every recovery model in this family becomes available: rebuilding a lost piece from its derivation, re-running a failed unit, or restoring accumulated state and rewinding to the matching position. ## The boundary moves; it does not disappear This is the point candidates miss, and the one worth saying out loud. The unrecoverable segment is now the hop between the feed and durable storage. Whatever that hop had received and not yet appended when it died is lost exactly as before — the segment is shorter, not absent. Therefore the design rule for that hop is **make it boring**: - No transformation, no enrichment, no joins, no lookups against anything that can be slow or down. - No parsing that can fail on the content of a record; append the bytes and let the real job interpret them. - Nothing that makes its restart slow, because its downtime is a window during which arriving records land nowhere at all. - Deploy and change it rarely, on its own schedule, separately from the processing logic that changes weekly. Every one of those rules exists for one reason: the hop's probability of failure and its time to restart are now the pipeline's loss budget. ## Narrowing the remaining loss What further reduction is possible depends entirely on what the feed offers, and this varies more than people expect: | What the feed offers | What the hop can do | Residual loss | |---|---|---| | Fire-and-forget delivery, no acknowledgement | append as fast as possible | everything received and not yet appended, plus everything sent while the hop is down | | Acknowledgement after receipt | write **a write-ahead record of intent** — durably record what it is about to do before doing it — then acknowledge only after the append is durable | only records in flight on the wire | | The sender retries on failure to acknowledge | as above, plus the sender re-sends | near zero, at the price of repeats the downstream must absorb | Note what the third row buys and costs: pushing the loss towards zero converts it into repetition, which then has to be made harmless further along. That is the same trade the ordering of a read position makes, one hop earlier. ## What it costs - **A hop of latency**, usually small, but real for a pipeline with a tight end-to-end budget. - **A system to run**: durable storage that has to be sized, monitored, paid for and kept alive, and whose own outage now stops the pipeline at the front rather than in the middle. - **Storage volume** proportional to how far back you want to be able to rewind, which is a deliberate decision rather than a default to accept. - **A second place a position is tracked**, and therefore a second place it can be wrong. - **Operational surface**: two components to deploy where there was one. ## Variants and what differs between runtimes Some runtimes ship a receiving step that can durably record what it has taken before the job acts on it, which is the same mechanism built in rather than deployed alongside. Others offer nothing of the kind, and the hop is yours to build and run. A third possibility, where the feed's own producer is under your control, is to change the producer to write durably first and let the job read that — the cheapest version of this design, and the one to ask about before building anything. The judgment an interviewer is looking for is not "add a buffer". It is that you can say **where the unrecoverable segment is after your change**, how large it is, and what it would take to shrink it again.

  • The receiving hop is down for two minutes. What happens to records sent during that time?
    They land nowhere unless the sender itself retains or retries them. That is why the hop's restart time is part of the loss budget and why it should carry no logic that slows a restart. If the sender can buffer or re-send on failure to acknowledge, the loss becomes repetition instead, which the pipeline can absorb downstream.
  • Could you skip the extra hop and have the processing job append the records itself?
    Only at the cost of coupling the durability of the input to the availability of the processing logic. Every deploy, every failure caused by a bad record and every slow operator becomes a gap in the input record itself. Separating them exists so that the part that must never stop is the part that does almost nothing.
  • How far back should the durable copy let you rewind?
    Far enough to cover the worst realistic gap between a failure and a working restart, including a bad deploy discovered the next morning. It is a recovery decision with a storage bill attached, not a value to leave at whatever the storage came with.

saying these in an interview costs you the question

  • Says adding a durable copy makes the pipeline fully recoverable.
  • Puts parsing and enrichment into the receiving hop to save a step.
  • Thinks restarting the process recovers records that were in flight.
  • Assumes the source's own delivery guarantee covers the job's recovery.
  • Treats how far back you can rewind as a setting rather than a decision.