skip to content

Batch and File-Based Ingestion

The unglamorous majority of real pipelines: scheduled extracts, files landing in object storage, and loads that must stay correct when they are re-run. Interviews here are less about tooling and more about idempotency, late data, and what happens when the source schema changes underneath you.

on this pageshow

explore

questions

18

What does an empty _SUCCESS marker file in a landing bucket tell a loader?

level: juniorimportance: must knowfreq 64%

answer

  1. object storage has no folder-level commit
  2. the reader needs to know when to start
  3. written last, deliberately
  4. its contents are empty; its presence is the signal

basics

~20 s

A _SUCCESS marker is written only after every data file in a batch has finished uploading, so its presence is the producer's promise that the drop is complete and safe to read. Its contents are irrelevant.

solid answer

~40 s

Object storage has no notion of a folder-level commit: each object becomes visible the moment its own upload finishes, so a loader that lists a prefix can easily see three of five files and load a partial batch. The convention is that the producer writes all data files first, then writes a zero-byte `_SUCCESS` (or `_COMMITTED`, `.done`) object **last**, and the consumer refuses to touch the prefix until that marker exists. The marker carries no data — its existence is the whole signal. It says "I finished writing what I intended to write"; it does not say the rows are valid, non-duplicated, or non-empty. Where you also need to know *which* objects belong to the drop, upgrade the marker to a manifest listing the object keys and row counts.

go deeper

for a junior

Recall the rule: data files first, empty marker last, and the loader waits for the marker before reading anything. Know that the marker's contents are empty and its presence is the entire message.

for a middle

Explain why per-object visibility makes a partial read likely without a marker, and where a manifest with file names and row counts is worth the extra producer work.

for a senior

Show the operational half: alert on a marker that never arrives, decide what a zero-row day looks like, and handle orphan files left by a failed run that the marker cannot see.

for a principal

Own it as a contract term with the producing team or partner — who writes the marker, when, what an empty day means, and what happens when the convention is broken. A convention nobody is accountable for degrades to noise.

## The problem a marker solves When the integration surface is "files in a bucket", the producer and consumer share no transaction. In a relational handoff the writer commits and the reader sees all-or-nothing. In object storage — S3, GCS, ADLS — there is no such boundary at the level of a prefix. Each object becomes visible independently, as soon as its own upload completes. A producer writing five Parquet parts creates five separate visibility events spread over however long the upload takes. A consumer that polls the prefix and loads whatever it finds will therefore, sooner or later, load two of five parts and declare the day loaded. Nothing errors. The row count is simply wrong, and it is wrong silently, which is the worst failure shape in ingestion. Note that this is not the old "eventual consistency" story. S3 has offered strong read-after-write consistency for new objects and for listings since late 2020, so a listing is no longer stale. The completeness problem survives that change, because uploads are per-object by nature: a correct listing of a half-finished drop is still a half-finished drop. ## The convention The pattern, inherited from Hadoop's output committers and now near-universal in lakes: 1. Producer writes every data file into the target prefix (or into a temporary prefix and copies them in). 2. Producer writes a zero-byte object named `_SUCCESS` last. 3. Consumer lists the prefix. If `_SUCCESS` is absent, it does nothing and tries again later. 4. Consumer loads the data files only when the marker is present. The file is empty on purpose. The information is entirely in its existence and in the ordering of the writes. Variants you will meet in the wild: `_COMMITTED`, `_DONE`, `.complete`, `<batch>.ok` sitting beside `<batch>.csv.gz`. The name does not matter; the write-it-last discipline does. ## What the marker does and does not guarantee It guarantees: the producer believes it finished. That is genuinely useful and it is all it is. It does **not** guarantee that the data is valid, that the schema is unchanged, that rows are unique, or that the batch is non-empty. It does not tell you whether this is the first or the fifth time this drop has appeared. And it does not tell you *which* objects belong to the batch — if a previous partial run left `part-0007.parquet` behind and the current run produced only six parts, listing the prefix picks up an orphan the marker knows nothing about. It also cannot help if the producer writes it in the wrong order, or writes it from a different process that merely assumes the upload finished. The marker is a convention, not an enforcement mechanism; a producer that writes it early has produced a lie the consumer will believe. ## Manifests: the stronger form Where the orphan-file or partial-rerun risk is real, replace the empty marker with a small manifest — typically JSON — that names every object in the drop and, ideally, carries a row count or byte count per file and a batch identifier. A manifest gives you three things a marker cannot. First, an exact file set: you load what the manifest names and ignore everything else in the prefix, so stragglers from a failed run are inert. Second, verifiability: after loading you can compare rows loaded against rows declared and fail loudly on a mismatch rather than shipping a short batch. Third, identity: a batch id makes it possible to recognise a re-drop and decide whether it replaces or supplements the previous one. The cost is that the producer must now emit it correctly, which is a contract negotiation rather than a code change when the producer is another team or an external partner. ## Operational notes Write the marker from the same process that wrote the data, after its final flush and close — not from a scheduler that assumes the job finished. If the producer stages files under a temporary prefix and copies them into place, write the marker after the copies, not after the staging writes. On the consumer side, treat a missing marker as *not yet ready*, not as an error, but alarm when a marker that was expected by a certain time has not appeared: a drop that never arrives and a drop that fails are both incidents, and only the second one usually pages anyone. Also decide explicitly what an empty day looks like — a marker with zero data files is a clear "nothing happened today", whereas an empty prefix is indistinguishable from a broken producer. Finally, the marker is orthogonal to idempotency. Knowing a drop is complete does not stop you loading it twice; that needs its own mechanism.

  • When is a zero-byte marker not enough, and what replaces it?
    When leftovers from a failed run can sit in the same prefix, or when you want to verify what you loaded. A manifest — JSON listing each object key plus a row or byte count and a batch id — lets the loader read exactly the named files, ignore orphans, and compare rows loaded against rows declared, failing loudly on a short batch instead of shipping it.
  • A producer writes _SUCCESS from its scheduler after the job exits. Why is that unsafe?
    The scheduler only knows the process ended, not that every buffer was flushed and every multipart upload completed. A job that dies after writing four of five parts can still exit in a way the scheduler reads as done. The marker must be written by the writer itself, after its final close, so it can only exist if the writes did.
  • How should a loader treat a prefix with data files but no marker for several hours?
    As not-ready, but alarm on it. Missing is not the same as failed: silently skipping forever means a day quietly never loads. Track expected drop times and page when a marker has not appeared by its deadline, so an absent producer is as visible as a crashed one.

It is the shipping label taped on last: the box may be full of the wrong parts, but at least you know nobody is still putting things in it.

saying these in an interview costs you the question

  • Assumes a whole prefix becomes visible atomically in object storage
  • Says the marker validates the data or the row counts
  • Writes the marker before or alongside the data files
  • Thinks waiting a few extra minutes reliably avoids partial reads
  • Treats a permanently missing marker as normal instead of alerting

context

open as a page

What is the difference between a full-refresh and an incremental extract from a source table?

level: juniorimportance: must knowfreq 78%

basics

~20 s

A full-refresh extract reads every source row each run and replaces the target. An incremental extract reads only rows changed since a stored watermark and merges them: far cheaper, but it must handle boundary rows, deletes and reruns itself.

open as a page

In a batch ingestion pipeline, what is schema drift and why does it break loads?

level: juniorimportance: must knowfreq 70%

basics

~10 s

Schema drift is the source changing shape without warning: columns added, dropped, renamed, retyped or reordered. Loads written against yesterday's shape then fail on a cast, shift positionally, or silently drop the new field.

open as a page

Why can an incremental extract filtered on updated_at above the stored watermark still lose rows?

level: middleimportance: must knowfreq 68%

basics

~20 s

Because updated_at is stamped when a row is written, not when its transaction commits, so a row can become visible only after the extract has already moved its watermark past that timestamp. Boundary ties and clock skew lose rows the same way.

open as a page

A file loader crashed mid-batch, re-ran, and now every row from that drop is duplicated. How do you make file ingestion idempotent?

level: seniorimportance: must knowfreq 66%

basics

~20 s

File ingestion is at-least-once by nature, so make replay harmless: track processed objects by key plus content hash in the same transaction as the load, or make the load replace a partition rather than append to it, or upsert on a deterministic row key.

open as a page

When should a batch load auto-evolve the target table's DDL instead of failing loudly?

level: seniorimportance: must knowfreq 62%

basics

~20 s

Auto-evolve additive, reversible changes in a raw landing zone — a new nullable column, a widened type — and always notify. Fail loudly on destructive or ambiguous ones: drops, renames, narrowing retypes, and anything in a published table consumers depend on.

open as a page

When a source can drop either CSV or Parquet into a landing bucket, what does Parquet buy you?

level: middleimportance: should knowfreq 56%

basics

~20 s

Parquet carries types and column names inside the file, so the loader stops guessing; it removes delimiter, quoting and encoding ambiguity, compresses well, and lets a reader fetch only the columns it needs. CSV's advantages are producer reach and human readability.

open as a page

Why do landing zones store dropped files under dated prefixes instead of one flat folder?

level: middleimportance: should knowfreq 66%

basics

~20 s

Dated prefixes keep listings bounded, give replay a natural unit you can reload or overwrite whole, let retention rules expire old data by prefix, and let query engines skip prefixes that cannot match a date filter.

open as a page

When should a batch pipeline advance its stored extraction watermark to a new value?

level: middleimportance: should knowfreq 54%

basics

~20 s

Only after the target write for that batch has durably committed, and to a value derived from the data actually loaded or the window's upper bound — never to the current wall-clock time, and never before the load succeeds.

open as a page

A source column changed from integer to free text overnight — how should the landing load react?

level: middleimportance: should knowfreq 50%

basics

~20 s

Widen the landing column to text so no value is lost, cast to the numeric type downstream, and route rows that fail the cast to a reject location. Never narrow the target back or drop the offending rows.

open as a page

A landing bucket gets thousands of 20 KB JSON files an hour and the hourly load keeps slowing. Why?

level: seniorimportance: should knowfreq 54%

basics

~20 s

Per-file fixed cost is dominating: paginated listing, a request and open per object, and one scheduled task per file. The bytes are trivial; the overhead is not. Fix it by batching at the producer or adding a compaction hop before the load.

open as a page

How do you make a re-run of an incremental batch load land the same window without duplicates?

level: seniorimportance: should knowfreq 62%

basics

~20 s

Give the run explicit half-open window bounds instead of reading whatever is new, then apply the slice with a keyed write: an upsert on the row's identity, or an atomic replacement of the partition the window covers. Both make a replay a no-op.

open as a page

Why do rows hard-deleted at the source never disappear from a watermark-based incremental extract?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Because a deleted row is gone from the table, so it can never satisfy a filter on a change column. The extract only ever sees rows that still exist, and the target keeps the stale copy forever with no error raised.

open as a page

A source dropped a column but your nightly load still succeeds with NULLs — how do you catch that?

level: seniorimportance: should knowfreq 48%

basics

~20 s

Compare the source's observed column list against the shape recorded from the last accepted run and fail on a missing column, and monitor per-column null rates so a column that goes from 2 percent null to 100 percent raises an alert on the first batch.

open as a page

What would you pin down in a partner's file-drop convention before accepting daily CSV drops?

level: principalimportance: should knowfreq 40%

basics

~20 s

Pin the delivery mechanics: a scoped prefix and deterministic file names, staged upload then copy to the final key, a completeness marker or manifest, explicit re-drop semantics, a mandatory zero-row drop on empty days, retention, and a named owner on each side.

open as a page

How would you backfill two years of history into a table that a nightly incremental load keeps current?

level: principalimportance: should knowfreq 40%

basics

~20 s

Run the backfill as a separate job with its own state, chunked into bounded half-open windows that each apply idempotently. Leave the nightly load running, throttle the source reads, and track chunk completion so failures resume instead of restarting.

open as a page

How would you design batch ingestion for schema drift from source systems you don't control?

level: principalimportance: should knowfreq 40%

basics

~20 s

Land permissively and immutably, publish strictly through explicit projections, and make drift detection a first-class pipeline step that classifies each change and routes it — auto-apply the safe ones with a notification, halt on the destructive ones, and tier sources by blast radius.

open as a page

A header-less CSV extract gains a column in position 3 — what happens to a positional load?

level: middleimportance: nice to knowfreq 42%

basics

~20 s

Every value after position 3 shifts one place, so the load writes each into the wrong target column. Where the shifted types happen to be compatible it succeeds and corrupts data silently; otherwise it fails on a cast.

open as a page