skip to content

File Landing Patterns

When the interface is "files in a bucket", the pipeline's correctness depends on conventions: how partitions are laid out, how you know a drop is complete, and how you avoid loading the same file twice. Small-file accumulation is the classic operational trap here.

on this pageshow

explore

questions

6

What does an empty _SUCCESS marker file in a landing bucket tell a loader?

level: juniorimportance: must knowfreq 64%

answer

  1. object storage has no folder-level commit
  2. the reader needs to know when to start
  3. written last, deliberately
  4. its contents are empty; its presence is the signal

basics

~20 s

A _SUCCESS marker is written only after every data file in a batch has finished uploading, so its presence is the producer's promise that the drop is complete and safe to read. Its contents are irrelevant.

solid answer

~40 s

Object storage has no notion of a folder-level commit: each object becomes visible the moment its own upload finishes, so a loader that lists a prefix can easily see three of five files and load a partial batch. The convention is that the producer writes all data files first, then writes a zero-byte `_SUCCESS` (or `_COMMITTED`, `.done`) object **last**, and the consumer refuses to touch the prefix until that marker exists. The marker carries no data — its existence is the whole signal. It says "I finished writing what I intended to write"; it does not say the rows are valid, non-duplicated, or non-empty. Where you also need to know *which* objects belong to the drop, upgrade the marker to a manifest listing the object keys and row counts.

go deeper

for a junior

Recall the rule: data files first, empty marker last, and the loader waits for the marker before reading anything. Know that the marker's contents are empty and its presence is the entire message.

for a middle

Explain why per-object visibility makes a partial read likely without a marker, and where a manifest with file names and row counts is worth the extra producer work.

for a senior

Show the operational half: alert on a marker that never arrives, decide what a zero-row day looks like, and handle orphan files left by a failed run that the marker cannot see.

for a principal

Own it as a contract term with the producing team or partner — who writes the marker, when, what an empty day means, and what happens when the convention is broken. A convention nobody is accountable for degrades to noise.

## The problem a marker solves When the integration surface is "files in a bucket", the producer and consumer share no transaction. In a relational handoff the writer commits and the reader sees all-or-nothing. In object storage — S3, GCS, ADLS — there is no such boundary at the level of a prefix. Each object becomes visible independently, as soon as its own upload completes. A producer writing five Parquet parts creates five separate visibility events spread over however long the upload takes. A consumer that polls the prefix and loads whatever it finds will therefore, sooner or later, load two of five parts and declare the day loaded. Nothing errors. The row count is simply wrong, and it is wrong silently, which is the worst failure shape in ingestion. Note that this is not the old "eventual consistency" story. S3 has offered strong read-after-write consistency for new objects and for listings since late 2020, so a listing is no longer stale. The completeness problem survives that change, because uploads are per-object by nature: a correct listing of a half-finished drop is still a half-finished drop. ## The convention The pattern, inherited from Hadoop's output committers and now near-universal in lakes: 1. Producer writes every data file into the target prefix (or into a temporary prefix and copies them in). 2. Producer writes a zero-byte object named `_SUCCESS` last. 3. Consumer lists the prefix. If `_SUCCESS` is absent, it does nothing and tries again later. 4. Consumer loads the data files only when the marker is present. The file is empty on purpose. The information is entirely in its existence and in the ordering of the writes. Variants you will meet in the wild: `_COMMITTED`, `_DONE`, `.complete`, `<batch>.ok` sitting beside `<batch>.csv.gz`. The name does not matter; the write-it-last discipline does. ## What the marker does and does not guarantee It guarantees: the producer believes it finished. That is genuinely useful and it is all it is. It does **not** guarantee that the data is valid, that the schema is unchanged, that rows are unique, or that the batch is non-empty. It does not tell you whether this is the first or the fifth time this drop has appeared. And it does not tell you *which* objects belong to the batch — if a previous partial run left `part-0007.parquet` behind and the current run produced only six parts, listing the prefix picks up an orphan the marker knows nothing about. It also cannot help if the producer writes it in the wrong order, or writes it from a different process that merely assumes the upload finished. The marker is a convention, not an enforcement mechanism; a producer that writes it early has produced a lie the consumer will believe. ## Manifests: the stronger form Where the orphan-file or partial-rerun risk is real, replace the empty marker with a small manifest — typically JSON — that names every object in the drop and, ideally, carries a row count or byte count per file and a batch identifier. A manifest gives you three things a marker cannot. First, an exact file set: you load what the manifest names and ignore everything else in the prefix, so stragglers from a failed run are inert. Second, verifiability: after loading you can compare rows loaded against rows declared and fail loudly on a mismatch rather than shipping a short batch. Third, identity: a batch id makes it possible to recognise a re-drop and decide whether it replaces or supplements the previous one. The cost is that the producer must now emit it correctly, which is a contract negotiation rather than a code change when the producer is another team or an external partner. ## Operational notes Write the marker from the same process that wrote the data, after its final flush and close — not from a scheduler that assumes the job finished. If the producer stages files under a temporary prefix and copies them into place, write the marker after the copies, not after the staging writes. On the consumer side, treat a missing marker as *not yet ready*, not as an error, but alarm when a marker that was expected by a certain time has not appeared: a drop that never arrives and a drop that fails are both incidents, and only the second one usually pages anyone. Also decide explicitly what an empty day looks like — a marker with zero data files is a clear "nothing happened today", whereas an empty prefix is indistinguishable from a broken producer. Finally, the marker is orthogonal to idempotency. Knowing a drop is complete does not stop you loading it twice; that needs its own mechanism.

  • When is a zero-byte marker not enough, and what replaces it?
    When leftovers from a failed run can sit in the same prefix, or when you want to verify what you loaded. A manifest — JSON listing each object key plus a row or byte count and a batch id — lets the loader read exactly the named files, ignore orphans, and compare rows loaded against rows declared, failing loudly on a short batch instead of shipping it.
  • A producer writes _SUCCESS from its scheduler after the job exits. Why is that unsafe?
    The scheduler only knows the process ended, not that every buffer was flushed and every multipart upload completed. A job that dies after writing four of five parts can still exit in a way the scheduler reads as done. The marker must be written by the writer itself, after its final close, so it can only exist if the writes did.
  • How should a loader treat a prefix with data files but no marker for several hours?
    As not-ready, but alarm on it. Missing is not the same as failed: silently skipping forever means a day quietly never loads. Track expected drop times and page when a marker has not appeared by its deadline, so an absent producer is as visible as a crashed one.

It is the shipping label taped on last: the box may be full of the wrong parts, but at least you know nobody is still putting things in it.

saying these in an interview costs you the question

  • Assumes a whole prefix becomes visible atomically in object storage
  • Says the marker validates the data or the row counts
  • Writes the marker before or alongside the data files
  • Thinks waiting a few extra minutes reliably avoids partial reads
  • Treats a permanently missing marker as normal instead of alerting

context

open as a page

A file loader crashed mid-batch, re-ran, and now every row from that drop is duplicated. How do you make file ingestion idempotent?

level: seniorimportance: must knowfreq 66%

basics

~20 s

File ingestion is at-least-once by nature, so make replay harmless: track processed objects by key plus content hash in the same transaction as the load, or make the load replace a partition rather than append to it, or upsert on a deterministic row key.

open as a page

When a source can drop either CSV or Parquet into a landing bucket, what does Parquet buy you?

level: middleimportance: should knowfreq 56%

basics

~20 s

Parquet carries types and column names inside the file, so the loader stops guessing; it removes delimiter, quoting and encoding ambiguity, compresses well, and lets a reader fetch only the columns it needs. CSV's advantages are producer reach and human readability.

open as a page

Why do landing zones store dropped files under dated prefixes instead of one flat folder?

level: middleimportance: should knowfreq 66%

basics

~20 s

Dated prefixes keep listings bounded, give replay a natural unit you can reload or overwrite whole, let retention rules expire old data by prefix, and let query engines skip prefixes that cannot match a date filter.

open as a page

A landing bucket gets thousands of 20 KB JSON files an hour and the hourly load keeps slowing. Why?

level: seniorimportance: should knowfreq 54%

basics

~20 s

Per-file fixed cost is dominating: paginated listing, a request and open per object, and one scheduled task per file. The bytes are trivial; the overhead is not. Fix it by batching at the producer or adding a compaction hop before the load.

open as a page

What would you pin down in a partner's file-drop convention before accepting daily CSV drops?

level: principalimportance: should knowfreq 40%

basics

~20 s

Pin the delivery mechanics: a scoped prefix and deterministic file names, staged upload then copy to the final key, a completeness marker or manifest, explicit re-drop semantics, a mandatory zero-row drop on empty days, retention, and a named owner on each side.

open as a page