skip to content

When a source can drop either CSV or Parquet into a landing bucket, what does Parquet buy you?

level: middleimportance: should knowfreq 56%

answer

  1. what does the loader have to guess
  2. one format ships its own schema
  3. leading zeros and embedded newlines
  4. the producer's reach versus the loader's safety

basics

~20 s

Parquet carries types and column names inside the file, so the loader stops guessing; it removes delimiter, quoting and encoding ambiguity, compresses well, and lets a reader fetch only the columns it needs. CSV's advantages are producer reach and human readability.

solid answer

~50 s

CSV is text with a dialect, and every dialect detail is a place to be wrong: delimiter, quote and escape characters, embedded newlines, encoding, header presence, null token, decimal separator, date format. Worse, types are inferred, so `01234` becomes `1234` and `1e3` becomes a float depending on who reads it. Parquet is self-describing and typed — the schema travels with the data, so what the producer wrote is what you load, and a new or renamed column is detectable rather than a silent column shift. It is column-oriented and compressed, so a reader can skip columns and move far fewer bytes. The costs: the producer needs a writer, you cannot eyeball or `grep` a file, and very small Parquet files carry proportionally more per-file overhead. If a partner can only send CSV, pin the dialect in writing, land it verbatim, and convert to Parquet in the first hop.

go deeper

for a junior

Know that a CSV file carries no types or dialect information while Parquet carries both, and that inferred types are where leading zeros and dates get mangled.

for a middle

List the dialect decisions a CSV drop leaves open and explain how a self-describing typed format removes them, along with the column-skipping and compression benefits and the readability cost.

for a senior

Argue the land-raw-then-normalise policy: keep the byte-exact original for replay and evidence, convert with an explicit schema, quarantine what does not fit.

for a principal

Weigh the negotiation cost against the engineering cost — pushing a format change onto a partner who gains nothing from it is a political project, and the honest answer is often to take their CSV and pin the dialect instead.

## What the format choice actually decides At the landing boundary the format is not an aesthetic preference; it decides how much the loader has to *guess*. Everything below follows from one difference: Parquet ships its schema with the data, and CSV does not. ## CSV: a dialect, not a format "CSV" names a family. Each drop implies decisions the file itself does not record: - delimiter (comma, tab, pipe, semicolon in locales that use comma decimals) - quote character and how a quote inside a quoted field is escaped (doubled, or backslashed) - whether fields may contain literal newlines inside quotes — which breaks any tool that splits on `\n` - character encoding: UTF-8 versus latin-1 versus UTF-16 with a BOM, a genuinely common partner problem - header row present or absent, and whether column order is stable - how null is represented: empty, `NULL`, `\N`, `NA`, `-` - timestamp format and timezone, decimal separator, thousands separator - line terminator, `\n` or `\r\n` Get any of these wrong and you do not get an error. You get shifted columns, truncated strings, or a whole file parsed as one row. On top of dialect sits **type inference**. A CSV field is a string until something decides otherwise. Readers that sniff types are the source of the classic damage: leading zeros stripped from zip codes and account numbers, `1e3` read as a float, a mostly-numeric column that turns into text the day a `N/A` appears, a phone number becoming scientific notation. And inference is *sample-based* in many tools, so the type can change between runs depending on what the first N rows happened to contain — the same file, loaded twice, yielding different types. ## Parquet: the schema travels Parquet files carry column names and types internally. The loader does not infer; it reads what the writer declared. That eliminates the entire inference class of bugs at a stroke, and it changes how you detect change: a column added at the source appears as a named field you can compare against expectation, rather than an extra positional value that silently shifts every column to its right. Being column-oriented and compressed also has direct operational effects. Files are typically much smaller than the equivalent CSV, which cuts storage and egress. A reader that needs three of eighty columns can fetch just those, so scans move far fewer bytes. And there is no dialect to agree on — nothing to misparse. ## What Parquet costs you The producer must be able to write it. Legacy systems, partner SFTP exports and "the finance team's weekly extract" often cannot, and demanding it is sometimes a six-month negotiation for a benefit the partner does not receive. It is opaque to ordinary tools. When a drop looks wrong at 2am, `head` and `grep` on a CSV answer the question in seconds; a Parquet file needs a reader. Keep one handy in the runbook. Small files are worse in Parquet than in CSV, proportionally: every file carries its own schema and footer metadata, so a stream of kilobyte-sized Parquet objects spends a large fraction of its bytes and its read latency on per-file overhead. Parquet rewards batching; if your producer emits micro-drops, fix the batching before the format. Finally, a compressed columnar file is less forgiving of partial corruption than a line-oriented text file where you can salvage the readable lines. ## Where JSON sits Newline-delimited JSON is the common middle ground: self-describing per record, tolerant of nested and irregular shapes, no dialect to negotiate, and still greppable. It costs a lot of bytes — every key repeated on every line — and it has weak typing (numbers, no distinction between integers and floats in many writers, dates as strings). It is a reasonable landing format for genuinely semi-structured sources, and a poor one for wide, regular, high-volume tables. ## The practical policy Take the best format the producer can actually emit, and land it **verbatim**. The landing zone's value is being a faithful archive of what was received; converting on the way in destroys your ability to prove what a partner sent and to reprocess with a fixed parser later. Then add a first normalisation hop that converts the raw drop to Parquet in a separate prefix, applying an explicit schema — not inference — and routing rows that do not fit to a quarantine location. Downstream loads read the normalised copy. This gives you the audit trail of the original and the type safety of the typed format, and it means an upstream dialect change is a parser fix and a replay rather than a corrupted table. When the drop must stay CSV, write the dialect down as part of the delivery agreement — delimiter, quote, escape, encoding, header, null token, timestamp format and timezone — and configure the reader explicitly against it rather than letting it sniff. Explicit beats clever every time on this boundary.

  • A partner can only export CSV. What do you put in the delivery agreement?
    The full dialect, in writing: delimiter, quote and escape characters, whether fields may contain newlines, character encoding, header row presence, the null token, decimal separator, line terminator, and the timestamp format with timezone. Then configure the reader explicitly against that dialect instead of letting it sniff, and land the file byte-for-byte before converting.
  • Why land the raw CSV at all if you immediately convert it to Parquet?
    Because the raw drop is your evidence and your replay source. If the parser was wrong — a mis-set encoding, a bad null token — you can fix it and reprocess from the original. Convert-on-arrival destroys that: the only surviving copy is the one you parsed incorrectly, and the partner may not be able to resend.
  • Why are very small Parquet files disproportionately wasteful?
    Each file carries its own schema and footer metadata plus a read round trip, so at kilobyte sizes the fixed per-file cost dominates the actual data. The columnar layout also has nothing to skip. The fix is upstream batching or a compaction hop, not a change of format.

saying these in an interview costs you the question

  • Lets the reader infer CSV types instead of declaring a schema
  • Assumes every CSV uses the same delimiter, quoting and encoding
  • Converts drops on arrival and keeps no raw copy
  • Claims Parquet is always better regardless of drop size
  • Ignores embedded newlines and quoted fields when splitting lines

context