skip to content

When would you choose csv.DictReader and csv.DictWriter over csv.reader and csv.writer?

level: middleimportance: should knowfreq 44%

answer

  1. Positional access versus access by name
  2. Column order stops mattering
  3. The header row becomes the keys
  4. restkey and restval handle ragged rows
  5. extrasaction raises by default on write

basics

~20 s

Use the dict-based pair when the file has a header row and you want fields addressed by column name rather than position, so a reordered column cannot break the code. Keep the list-based pair for headerless files and the hottest loops.

solid answer

~50 s

`csv.DictReader` reads the first record as the header - unless you pass `fieldnames=` - and yields each subsequent record as a plain `dict` keyed by column name. That decouples your code from column order, which is the main reason to prefer it: a producer who inserts a column in the middle breaks positional indexing and does not break name-based access. It also handles ragged rows explicitly: `restval=` supplies a value for missing columns and `restkey=` collects surplus fields into a list under a key you choose. `csv.DictWriter` is the mirror: you must give it `fieldnames=`, call `writeheader()` yourself, and by default it raises `ValueError` if a row dict contains a key you did not declare - `extrasaction="ignore"` relaxes that. The cost is a dict allocation per record, so for very large files or hot loops `csv.reader` is meaningfully cheaper.

code

python · 12 lines
python
import csv
import io

text = "id,lat,lon\n7,48.85,2.35\n8,52.52,13.40\n"
rows = list(csv.DictReader(io.StringIO(text, newline="")))
print(rows[0]["lat"], type(rows[0]))

out = io.StringIO(newline="")
w = csv.DictWriter(out, fieldnames=["id", "lat"], extrasaction="ignore")
w.writeheader()
w.writerows(rows)
print(out.getvalue().splitlines())

go deeper

for a junior

Know that DictReader turns the header row into dictionary keys so you write row['city'] instead of row[3]. Remember that DictWriter needs fieldnames up front and that you call writeheader() yourself.

for a middle

Explain the tradeoffs: name-based access survives a reordered column, at the cost of a dict per record. Be able to describe restval, restkey and the extrasaction default of raise, and what each one protects you from.

for a senior

Show that you validate fieldnames against the columns you require before processing, and that you handle header noise - stray whitespace, a byte-order mark, inconsistent case - rather than letting a KeyError surface deep into a long run.

for a principal

Own the contract between producer and consumer. Decide whether the header is a schema you enforce, and where validation and typing belong, given that the csv module guarantees record shape and nothing about meaning.

### The choice in one sentence `csv.reader` gives you positional records; `csv.DictReader` gives you named ones. Almost every real file has a header row, and almost every real consumer wants two or three named columns out of twenty, so the dict pair is the default choice and the list pair is the optimisation. ### DictReader mechanics `csv.DictReader(f)` consumes the first record of the file as the header and exposes it as the `fieldnames` attribute. Iterating yields one `dict` per subsequent record, mapping each header name to that record's `str` value. If the file has no header, pass `fieldnames=[...]` explicitly and the first record is then treated as data. Reading `fieldnames` before iteration is what triggers the header read, so it is safe to inspect it up front - useful when you want to fail fast on a file whose columns are not what you expected. Two arguments deal with the fact that real files are ragged: - **`restval=`** is the value used for declared columns a short record does not supply. Without it, missing keys map to `None`, which is a different failure mode from the empty string you get for a present-but-empty field. That distinction matters: `None` means the column was absent, `''` means it was there and blank. - **`restkey=`** names a key under which surplus fields of a long record are collected, as a list. Without it, the extra values are silently discarded. Since Python 3.8, `csv.DictReader` yields ordinary `dict` objects; earlier versions produced `collections.OrderedDict`. Both preserve insertion order, so the change is mostly invisible, but an equality assertion written against `OrderedDict` in an old test will notice. ### DictWriter mechanics `csv.DictWriter(f, fieldnames=[...])` is stricter than its reader, and deliberately so. `fieldnames` is mandatory, because the writer has to decide a column order and it will not guess one from the first dict it happens to see. The header is not written automatically; `writeheader()` is a separate call, so that appending to an existing file does not duplicate it. The strictness that surprises people is `extrasaction`. Its default, `"raise"`, means that passing a row dict with a key outside `fieldnames` raises `ValueError`. That is a feature: it catches the case where an upstream stage started emitting a new field and your output would otherwise have dropped it without a word. When you genuinely want a projection - carrying a wide dict through and writing three columns of it - set `extrasaction="ignore"` and be explicit about the intent. A *missing* key is handled by `restval=` instead, which defaults to the empty string. ### When to stay with the list pair Three situations argue for `csv.reader` and `csv.writer`: 1. **No header.** A file whose columns are defined by an external contract has nothing to key on, and inventing names is not always clearer than indexing. 2. **Volume.** `csv.DictReader` builds a dict per record. Over tens of millions of records that allocation is measurable, and if your loop reads two columns out of thirty, positional access with a couple of indices from the header is faster. Measure before you contort the code. 3. **Genuinely positional work** - transposing, re-emitting a file unchanged, or streaming rows through untouched, where naming the columns adds nothing. ### The habit worth forming Both pairs share the same dialect machinery, the same `newline=''` requirement on the file object, and the same guarantee that every value is a `str`. Choosing between them is a readability and robustness decision, not a correctness one - neither can save you from a column that changed meaning rather than position. One trap deserves naming. Because `DictReader` is keyed by header text, it inherits whatever the header actually says: trailing spaces, a byte-order mark on the first column of a file written by a spreadsheet, an inconsistent case. `row['id']` then raises `KeyError` on a file that looks correct in an editor. Normalising `fieldnames` right after construction - stripping and lower-casing - is a cheap defence, and reading the encoding correctly is what removes the byte-order mark. It is also why an explicit up-front check of `fieldnames` against the columns you require beats discovering the mismatch three million records into a run. ### Reading a headerless file with names anyway The two pairs are not a fork in the road you can only take once. A file with no header can still be read by name: pass `fieldnames=['id', 'lat', 'lon']` to `csv.DictReader` and the first record is treated as data rather than consumed as a header. That is often the better choice when column meanings come from an external contract, because the names then live in one visible place in your code instead of being scattered as bare indices through the loop. The mirror case also holds: `csv.DictWriter` will happily emit a headerless file if you simply never call `writeheader()`, which is what you want when appending to a file that already has one. Neither pair does anything about a column whose *meaning* changed while its name stayed the same, and no amount of dict access will catch that. Name-based access buys robustness against reordering and insertion, which is the common producer-side change; anything beyond that is validation you have to write.

  • What happens if a row dict passed to csv.DictWriter.writerow has a key that is not in fieldnames?
    It raises `ValueError`, because `extrasaction` defaults to `"raise"`. That is usually what you want - it catches an upstream stage that started emitting a field your output would otherwise drop silently. Pass `extrasaction="ignore"` when you deliberately want to project a wide dict down to a few declared columns.
  • How does csv.DictReader behave on a record with fewer fields than the header?
    Missing columns take the value of `restval`, which defaults to `None`. That is deliberately distinguishable from a present-but-empty field, which comes back as `''`. A record with more fields than the header puts the surplus in a list under the `restkey` key, and discards them entirely if you did not supply one.
  • Is there a reason to prefer csv.reader for a very large file?
    Yes - `csv.DictReader` allocates a dict per record, and across tens of millions of records that shows up in both time and garbage-collection pressure. If the loop touches two of thirty columns, reading the header once and indexing positionally is cheaper. It is a measured optimisation, not a default; the readability of named access usually wins.

saying these in an interview costs you the question

  • Thinks DictWriter writes the header row automatically
  • Assumes DictReader converts numeric columns
  • Says column order still matters with DictReader
  • Cannot say what happens to surplus fields in a long record
  • Believes DictWriter infers fieldnames from the first row

context