A job's input is a directory of files rather than a positioned log. What must hold for a restart to resume it correctly?
answer
- no single position over a directory
- immutable once visible
- whole file or no file
- the position is a set of identities
- path alone is not an identity
basics
~20 sFiles must be immutable once visible, become visible in one atomic step, and survive until the job can no longer need them. The recorded position is then a set of file identities, and each identity must change whenever the content does.
solid answer
~50 sA positioned log gives a reader a single number to resume at. A directory gives nothing of the kind, so re-readability has to be constructed out of the filesystem's behaviour. Three properties carry it. **Immutability**: a file already visible is never rewritten or appended in place, so a second read returns the same bytes and recomputed output agrees with what the crashed run produced. **Atomic appearance**: a producer writes under a temporary name and promotes it in one step, so a reader can never pick up a half-written file whose later re-read differs. **Retention**: nothing deletes a file the job might still have to re-read, which rules out deleting on successful read. The position itself becomes a set of file identities rather than a number, and an identity must be more than a path — a path reused for new content is invisible to a record that only remembers names.
go deeper
Recall the three conditions for a file input to be resumable: files never change after they appear, they appear whole rather than partially, and they are not deleted once read.
Explain why a directory has no single position, so progress is a set of file identities, and why that identity has to change whenever the content does.
Show that you check the producer's write pattern and the storage's listing behaviour before promising resume semantics, and that you can name the symptom — two runs disagreeing — that a mutable input produces.
Treat it as a contract with whoever produces the files: write-once, promote atomically, never reuse a name, retain long enough to reprocess, and price the storage that last clause requires.
## What a positioned log gives you, and a directory does not A source with positions hands a reader one compact thing to remember: *I have consumed up to here*. Resuming is then unambiguous, and the source itself enforces that the records behind that position do not change. A directory of files offers none of it. There is no total order, no single marker, and no rule stopping anyone from rewriting a file you already read. Re-readability over files is therefore not a given — it is a set of guarantees that the producers of those files, and the retention policy over them, have to actually provide. When they do, a file input is one of the most comfortable inputs to recover from; when they do not, the failure is quiet and shows up as two runs disagreeing. ## The three properties a restart depends on 1. **Files are immutable once visible.** Recovery is reprocessing, and reprocessing is only correct if a second read at the same place returns the same records. A file that is appended to, or rewritten with corrected content, breaks that: the re-run computes a piece from bytes that differ from the ones the crashed run used, and one output ends up assembled from two versions of the input. New data must arrive as new files, and a correction must arrive as a new file with a new identity — never as an edit of one already read. 2. **A file becomes visible in a single act.** Writers stage to a temporary name and promote with one atomic step, so the directory never shows a file that is still being written. Without this, a reader can consume half a file, and a re-read after the writer finishes returns more than the first read did. This is the same shape as **staged output promoted in a single step** — each unit writes privately and one final act makes the whole result visible — applied to the producer of your input rather than to your own output. 3. **Nothing removes a file the job might still need.** Deleting a file on successful read turns a re-readable source back into one that forgets: the moment a downstream failure forces reprocessing, the input for the affected range is gone. Move-on-read has the same effect unless the destination of the move is itself part of what a restart will look at. ## Recording a position that is a set Because there is no single number, the **recorded read position** — the durable marker of how far the job has got — becomes a record of *which files have been fully accounted for*. Two consequences: - **Identity must track content.** A record of paths alone cannot tell that a path was reused for different content; a restart then skips data that was never processed. Pair the path with something that changes when the content does — a size and modification time, a content digest, or a naming scheme that forbids reuse outright. - **The ordering rule still governs.** A file is marked accounted for only after the output derived from it is durable. Marking on read reproduces the silent-loss failure exactly: the restart skips a file whose results never landed. ## Where runtimes differ This is one of the places the three lineages genuinely diverge, so avoid stating any one behaviour as the rule: | Approach | How a restart knows what to skip | Where it strains | |---|---|---| | A finite job simply re-runs entirely | nothing is skipped; the whole directory is read again and the result is republished in one promotion | wasteful on large inputs; wrong if the directory has grown since the failed run began | | The job keeps consumed-file identities in its own durable state | the state names them, and the resume lists the directory and subtracts | that state must be written under the ordering rule, and it grows with the file count | | Progress is re-derived from what the destination already holds | the output itself records the range it covers | needs a destination that can be queried that way, and per-file granularity is often lost | A continuous job watching a directory for new arrivals has an extra hazard: listing is not instantaneous and not always consistent, so a file that appears with a timestamp behind the watermark of what has already been listed can be missed entirely. That is a property of the storage system's listing behaviour, not of the engine, and it is worth asking about before promising exactly what a restart will pick up. ## The failure you will actually see None of this announces itself. The symptoms are a re-run producing a total that differs from the original, or a day's data missing after a restart that reported success. Both trace back to the same root: the input was assumed to be re-readable because the files were still sitting there, when re-readable actually means *unchanged, whole, still present, and identifiable*.
- A producer appends new rows to yesterday's file each hour. What does that do to recovery?It removes re-readability. A re-read of that file returns more rows than the original read did, so recomputed output disagrees with what the crashed run produced, and the same data may be counted twice. Hourly arrivals have to be new files; the existing one must never change once a reader can see it.
- Why is recording only file paths as the position dangerous?Because a path can be reused with different content. The restart sees a name it has already processed and skips it, losing that data silently. Pair the path with a content digest, or with size and modification time, or forbid reuse by naming files so that each is written once.
- Is re-running the whole directory a legitimate recovery strategy?For a bounded job it often is, provided the output is staged and promoted in one act so the previous partial result never becomes visible, and provided the directory's contents are fixed for the run. It stops being viable when the input is large enough that the re-run misses its deadline, or when the directory keeps growing underneath.
saying these in an interview costs you the question
- Assumes files sitting in storage are automatically re-readable input.
- Believes appending to a file a reader has already seen is harmless.
- Deletes each input file as soon as it has been read successfully.
- Records processed file paths without anything that changes with the content.
- Marks a file as processed at read time rather than after its output is durable.