Why is listing a directory of Parquet files an unreliable way to define a table's contents?
answer
- storage has no notion of a transaction
- a job's output appears file by file
- what a reader sees mid-rewrite
- the folder is a snapshot of bytes, not commits
basics
~20 sA directory listing shows whatever is on storage at that instant, including half-written job output and files a rewrite is about to delete. A table format instead reads an explicit versioned list of files, so every reader sees one consistent set.
solid answer
~50 sStorage is not transactional, so a listing is a snapshot of *bytes that exist*, not of *rows that are committed*. A job writing 500 files makes its output visible file by file, so a query running concurrently reads a partial result. A rewrite that replaces files has a window where both old and new files are present — duplicated rows — or where a reader's file has already been deleted, so the query fails. Nothing enforces that every file in the directory shares the table's schema. And on object storage a listing is a paginated, per-prefix scan whose cost grows with file count, so planning a table with millions of files is slow before a single byte of data is read. A table format replaces "whatever is in the folder" with an explicit list of files per version, swapped atomically.
code
text · 7 lines# Directory listed while a 500-file write is in flight
s3://lake/orders/dt=2026-03-01/
part-00000.parquet # committed by the running job
part-00001.parquet
part-00002.parquet
... # 497 files not written yet
# A query listing now returns a partial table -- successfully, with no errorgo deeper
Recall the simplest case: a job is halfway through writing its files and a query runs. With a plain directory that query returns partial data and no error. That single scenario is the whole argument in miniature.
Explain all four mechanics — no atomic multi-file write, no isolation during rewrite, no authoritative schema, and expensive listing — and describe how an explicit versioned file list replaces each one.
Bring the operational scars: duplicated rows during overwrite windows, file-not-found failures in long queries, hour-long planning on wide partitioned tables, and metastore partition drift needing repair jobs.
Frame it as a correctness guarantee the platform provides once, centrally, rather than something every pipeline author reimplements with staging paths and success markers — and price the migration against the cost of those recurring incidents.
## The Hive-style assumption The pre-lakehouse arrangement was simple: a metastore entry mapped a table name to a directory, and, for a partitioned table, each partition value to a subdirectory. The table's contents were *defined by* what was in those directories at read time. That works fine on a quiet HDFS cluster with one batch job a night. It breaks in every other situation, and each break is a favourite interview question. ## Failure 1: no atomicity A writer producing 500 output files creates them one at a time. Storage has no notion of "these 500 appear together". Any query that lists the directory in the middle of that job sees an arbitrary prefix of the output — a partial, wrong answer that returns successfully and silently. There is no error to alert anyone. Table formats fix this by writing all the data files first and then making a single atomic metadata commit that moves the table's current version from the old file set to the new one. Until that instant, readers see the old set; after it, the new set. Never a mixture. ## Failure 2: no isolation from rewrites An overwrite or compaction has to both add new files and delete old ones. With a directory as the source of truth, there is an unavoidable window where both are present — every affected row appears twice — or where the old file has been deleted while a running query still holds its path, so the query dies with a file-not-found error. Some engines papered over this with directory rename tricks, but object stores like S3 have no atomic directory rename; a "rename" is a copy of every object followed by deletes, which is neither atomic nor cheap. In a table format the same rewrite is one commit that adds the new files and removes the old ones together, and older versions keep referencing the old files until a retention-based cleanup job removes them, so in-flight readers finish safely. ## Failure 3: no schema authority Each Parquet file carries its own schema. If two jobs write with slightly different schemas — a renamed column, an int that became a long, a nullable field — the directory happily contains both, and the reader has to reconcile them at query time or fail. There is no place that says what the table's schema *is*. The table layer holds one canonical schema and defines how files written under older schemas map onto it, which is what makes safe column add, drop and rename possible without rewriting data. ## Failure 4: listing cost and consistency On object storage, listing is a paginated API call per prefix, returning a limited number of keys per request. A table with a million files needs thousands of round trips before planning can even begin, and a deeply partitioned table multiplies that by the number of prefixes. This is the notorious slow-planning problem of large Hive tables. Historically S3 listings were also eventually consistent, so a just-written file could be absent from a listing entirely; Amazon has since made them strongly consistent, but the cost problem remains. A table format reads its file list from a small number of metadata objects instead, and that list already carries per-file statistics, so many files are eliminated from the plan without being listed or opened at all. ## Failure 5: no history and no way back With a directory, the previous state of the table is simply gone once files are deleted. There is no rollback for a bad overwrite and no way to reproduce yesterday's report. Versioned metadata gives both for free, because old versions reference file sets that are retained until expiry. ## What replaces it The table layer keeps, for each version, an explicit list of data files with their statistics, and a single pointer to the current version. Commits are optimistic: a writer reads the current version, prepares its files, and then swaps the pointer only if nothing else has committed in the meantime; otherwise it retries. That one indirection — a pointer to a file list, rather than a directory scan — is what turns a pile of Parquet into a table. ## The nuance worth stating Directory listing is not merely slower; it is *incorrect* under concurrency. A candidate who only says "listing is slow on S3" has spotted the performance symptom and missed the correctness argument, which is the real reason table formats exist.
- Why did engines historically write to a staging path and then rename the directory?Renaming a completed output directory into place was a cheap way to approximate an atomic commit on HDFS, where directory rename is a single metadata operation. It never worked properly on object storage: S3 has no atomic directory rename, so the operation becomes a copy of every object plus deletes — slow, non-atomic, and interruptible halfway. This limitation is a large part of why table formats appeared.
- Does a partitioned directory layout fix any of these problems?No. It narrows how many prefixes a query must list, which helps pruning, but each partition directory is still just a folder with the same atomicity, isolation and schema-authority problems. It also adds its own failure: the metastore's partition list can drift out of sync with what is actually on storage, so repair commands become part of routine operations.
- How does a table format let a long-running query survive a concurrent rewrite?The reader pins the table version it started with and resolves its file list from that version's metadata. A concurrent commit creates a new version but does not touch the files the old version references. Files only disappear when a retention-based cleanup job removes versions older than the retention window — which is why that window must exceed your longest query.
saying these in an interview costs you the question
- Saying the only problem with directory listing is that it is slow
- Assuming a rename makes a multi-file write atomic on S3
- Believing readers and writers never overlap in practice
- Thinking per-file schemas add up to a table schema
- Claiming partitioned directories solve concurrency