skip to content

In Ruby 4.0, what does CSV.foreach(path, headers: true) yield for each line, and what do the converters and header_converters options change?

level: middleimportance: should knowfreq 45%

answer

  1. a bundled gem since 3.4
  2. CSV::Row, fields are Strings
  3. row["price"], nil if absent
  4. :numeric uses Integer() and Float()
  5. header_converters: :symbol

basics

~20 s

With headers: true, CSV.foreach yields one CSV::Row per data line, indexed by header name, with every field a String. converters: :numeric turns numeric-looking fields into Integer or Float; header_converters: :symbol turns headers into snake_case Symbols.

solid answer

~40 s

`csv` is a **bundled gem** since Ruby 3.4, so under Bundler `require "csv"` raises `LoadError` until the Gemfile lists it. `CSV.foreach(path, headers: true)` streams the file one line at a time: the first line becomes the headers and each later line arrives as a `CSV::Row`, so `row["price"]` returns that column's value, always a `String` unless converted, and `nil` for an unknown header (`row.fetch("price")` raises `KeyError`). `converters: :numeric` runs `Integer()` then `Float()` on each field and keeps the String if both fail, which also means `"0100"` becomes `64`, because `Integer()` reads a leading zero as octal. `header_converters: :symbol` downcases headers, strips punctuation and joins words with underscores, so `"Unit Price"` becomes `:unit_price`. `CSV.table(path)` applies all three options at once but reads the whole file into memory.

code

ruby · 18 lines
ruby
require "csv"

# prices.csv:
# SKU,Name,Unit Price
# 0100,Desk lamp,24.50

CSV.foreach("prices.csv", headers: true) do |row|
  row["SKU"]         # => "0100"
  row["Unit Price"]  # => "24.50"
  row["Price"]       # => nil (no such header)
end

cents = ->(field, info) { info.header == :unit_price ? (field.to_r * 100).to_i : field }

CSV.foreach("prices.csv", headers: true, header_converters: :symbol, converters: [cents]) do |row|
  row[:sku]          # => "0100"  still a String
  row[:unit_price]   # => 2450
end

go deeper

for a junior

Recall that headers: true yields CSV::Row objects indexed by header name, and that every field starts as a String.

for a middle

Explain foreach versus read, the :numeric and :symbol converters, nil versus fetch for unknown headers, and the bundled-gem LoadError.

for a senior

Protect identifier and money columns with column-aware converters, stream large files, and handle byte-order marks in supplier exports.

for a principal

Standardise an import pipeline: declared column types, validation per row, reported line numbers, and no blanket converters.

## Loading the library Since Ruby 3.4, `csv` is a **bundled gem** rather than a default gem: it is installed with Ruby but is not activated automatically inside a Bundler-managed application. Ruby 4.0.7 bundles csv 3.3.5. Under Bundler, `require "csv"` without a Gemfile entry raises `LoadError` with a message saying csv "is not part of the default gems since Ruby 3.4.0" and suggesting adding it to the Gemfile. Outside Bundler it loads as before. ## Rows with headers Without options, CSV yields each line as an `Array` of Strings. With **`headers: true`** the first line is taken as the header row, and every later line is yielded as a **`CSV::Row`**: - `row["sku"]` (an alias of `Row#field`) returns the field under that header, or **`nil`** when no such header exists, so a typo is silent; - `row.fetch("sku")` raises `KeyError` ("key not found: sku") for an unknown header, and accepts a default or a block; - `row.to_h` returns a Hash from header to value; - fields are **Strings** (or `nil` for an empty unquoted field) until you convert them. ## Streaming versus loading | Call | Reads | Returns | |---|---|---| | `CSV.foreach(path, headers: true) { \|row\| }` | one line at a time | yields `CSV::Row`s; an `Enumerator` without a block | | `CSV.read(path, headers: true)` | the whole file | a `CSV::Table` | | `CSV.parse(string, headers: true)` | a String already in memory | a `CSV::Table`, or yields rows with a block | | `CSV.table(path)` | the whole file | a `CSV::Table` with `headers: true`, `converters: :numeric`, `header_converters: :symbol` | For a large supplier file, `CSV.foreach` keeps memory flat; `CSV.read` and `CSV.table` hold every row at once. ## Field converters The `converters:` option transforms each field after parsing: 1. `:integer` calls `Integer(field)` and keeps the String if that raises. 2. `:float` does the same with `Float(field)`. 3. `:numeric` is `[:integer, :float]`: try Integer, then Float. 4. `:date` and `:date_time` parse recognised date formats; `:all` is `:date_time` plus `:numeric`. 5. A **lambda** is a custom converter. With one parameter it receives the field; with two it also receives a `CSV::FieldInfo` whose `header`, `index` and `line` tell you which column you are converting. Two traps sit inside `:numeric` for a product catalogue: - **Leading zeros**: `Integer("0100")` is `64`, because `Kernel#Integer` treats a leading `0` as an octal prefix, and `"0109"` fails as octal and survives as a String. An SKU or postcode column must never be run through `:numeric`. - **Money as Float**: `"19.99"` becomes the Float `19.99`, with binary rounding. A custom converter that turns the price column into integer cents, or into `BigDecimal` (also a bundled gem since 3.4), keeps it exact. Apply converters to specific columns with a two-argument lambda that checks `info.header`, rather than converting every column blindly. ## Header converters `header_converters:` rewrites the header names once: - `:symbol` downcases, removes characters other than word characters and spaces, strips, and replaces runs of spaces with `_`, then makes a Symbol: `"Unit Price (EUR)"` becomes `:unit_price_eur`; - `:symbol_raw` makes a Symbol without any cleaning; - `:downcase` only downcases. After `:symbol`, rows must be indexed with Symbols: `row[:unit_price_eur]`, and `row["Unit Price (EUR)"]` returns `nil`. ## Writing CSV back out The same library writes files. `CSV.open(path, "w", write_headers: true, headers: ["sku", "price_cents"])` opens a file for writing and emits the header row first; `csv << row` then appends each row, given as an Array in header order or as a `CSV::Row`. `CSV.generate { |csv| csv << [...] }` builds the same output in a String, and `CSV.generate_line(["LMP-01", 2450])` formats a single line. Quoting of commas, quotes and newlines inside fields is handled for you on the way out. ## A byte-order mark Spreadsheet exports often start with a UTF-8 **byte-order mark**. CSV does not strip it, so the first header becomes `"sku"` and `row["sku"]` is `nil` on every row. Open the file with `encoding: "bom|utf-8"` (`CSV.foreach(path, headers: true, encoding: "bom|utf-8")`) to remove it.

  • Why does an SKU of 0100 become 64 with converters: :numeric?
    The `:integer` converter calls `Kernel#Integer`, which honours radix prefixes: a leading `0` means octal, so `"0100"` parses as 64, while `"0109"` is not valid octal and stays a String. Identifier columns should never pass through `:numeric`; convert only the columns that are really numbers.
  • When would you choose CSV.foreach over CSV.read?
    For large files. `CSV.foreach` opens the file and yields one row at a time, so memory stays flat, while `CSV.read` and `CSV.table` build a `CSV::Table` holding every row. `CSV.read` fits small files you need to traverse several times.
  • Why might row["sku"] be nil on every row of an exported spreadsheet?
    The file probably starts with a UTF-8 byte-order mark, which CSV keeps as part of the first header, so the header is not exactly `"sku"`. Opening with `encoding: "bom|utf-8"` strips it; `row.headers` shows the real header strings.

saying these in an interview costs you the question

  • require "csv" always works because csv is part of the standard library
  • headers: true makes CSV convert numeric fields automatically
  • row["missing"] raises KeyError for an unknown header
  • converters: :numeric is safe for identifier columns
  • CSV.foreach loads the whole file into memory first