How do you configure a FlatFileItemReader to parse a delimited CSV file into domain objects?
answer
- Resource → LineMapper → FieldSet → object
- DelimitedLineTokenizer + BeanWrapperFieldSetMapper
- linesToSkip for header
- FlatFileParseException → skippable
- implements ItemStream (line count)
basics
~20 sFlatFileItemReader reads a text file line by line. You give it a Resource, a LineMapper that splits each line (a DelimitedLineTokenizer into a FieldSet) and maps the fields to an object (a FieldSetMapper), plus optional linesToSkip for headers.
solid answer
~40 sFlatFileItemReader<T> reads a text Resource one line at a time and turns each line into a T via a LineMapper. The default DefaultLineMapper composes two collaborators: a LineTokenizer (usually DelimitedLineTokenizer for CSV or FixedLengthTokenizer for fixed-width) that parses a raw line into a FieldSet (indexed/named fields), and a FieldSetMapper (e.g. BeanWrapperFieldSetMapper, or a custom one) that maps the FieldSet to your domain object. You typically set the resource, column names, linesToSkip (to skip a header), encoding, and strict mode. In practice most people use FlatFileItemReaderBuilder. Parse failures throw FlatFileParseException, which you can skip in a fault-tolerant step. It implements ItemStream, tracking the current line number in the ExecutionContext so a restart resumes mid-file.
code
java · 14 lines@Bean
public FlatFileItemReader<Customer> customerReader() {
return new FlatFileItemReaderBuilder<Customer>()
.name("customerReader") // used as ExecutionContext key prefix
.resource(new ClassPathResource("customers.csv"))
.linesToSkip(1) // skip header row
.encoding("UTF-8")
.delimited()
.delimiter(",")
.names("id", "firstName", "lastName", "email")
.targetType(Customer.class) // BeanWrapperFieldSetMapper under the hood
.strict(true)
.build();
}go deeper
Knows it reads a CSV line by line into objects.
Can wire tokenizer + FieldSetMapper and skip headers via the builder.
Adds fault-tolerant skip, encoding, strict, custom FieldSetMapper, restartability.
Weighs flat-file streaming/restart vs. thread-safety and skip/audit policy at scale.
### What it is `FlatFileItemReader<T>` is the standard reader for **text files** — CSV, tab-separated, or fixed-width. It reads a Spring `Resource` **line by line** and converts each line into a domain object of type `T`. ### The pipeline: line → FieldSet → object `FlatFileItemReader` delegates conversion to a `LineMapper<T>`. The default is `DefaultLineMapper<T>`, which wires two pieces: 1. **`LineTokenizer`** — splits a raw `String` line into a **`FieldSet`** (an array of fields accessible by index or by name, with typed getters like `readString`, `readInt`, `readDate`). - `DelimitedLineTokenizer` — delimiter-based (default comma); handles quoted fields. - `FixedLengthTokenizer` — column ranges for fixed-width files. 2. **`FieldSetMapper<T>`** — maps the `FieldSet` to an object. - `BeanWrapperFieldSetMapper` — maps field **names** to bean properties via setters. - A custom `FieldSetMapper` for full control (type conversion, validation). ### Typical configuration knobs - `resource(...)` — the input file (`FileSystemResource`, `ClassPathResource`). - `linesToSkip(1)` — skip a header row (or use `skippedLinesCallback`). - `delimited().names("a","b")` / `fixedLength().columns(...)`. - `encoding(...)` — default is the platform default; set it explicitly. - `strict(true)` — throw if the resource is missing (set `false` to tolerate absent optional files). - `comments("#")` — ignore comment lines. - `targetType(MyDto.class)` (builder) — shortcut for `BeanWrapperFieldSetMapper`. ### The builder (idiomatic) Use `FlatFileItemReaderBuilder` rather than wiring `DefaultLineMapper` by hand. ### Restartability `FlatFileItemReader` implements `ItemStream`. In `open()` it opens the file (and skips forward to the saved line on restart); in `update()` it stores the **current line count** in the `ExecutionContext`; in `close()` it releases the file handle. So if a job fails at line 5,000 of 10,000, a restart resumes near line 5,000 rather than re-reading everything — provided `saveState` is left `true` (the default). ### Gotchas & edge cases - A malformed line throws `FlatFileParseException` (wrapping `IncorrectTokenCountException`, etc.). Add `.faultTolerant().skip(FlatFileParseException.class)` to tolerate bad rows. - **Not thread-safe** — wrap in `SynchronizedItemStreamReader` for multi-threaded steps. - Forgetting `linesToSkip(1)` makes the header row parse as data. - Wrong `encoding` silently corrupts non-ASCII data. - `strict(true)` (default) fails the step if the file is missing — sometimes you want `false`. - `BeanWrapperFieldSetMapper` requires the tokenizer's field **names** to match bean property names; a mismatch yields nulls, not errors.
- What is the role of a FieldSet, and how does it differ from a LineTokenizer?The LineTokenizer parses a raw line into a FieldSet — a structured, index/name-addressable collection of fields with typed accessors (readInt, readDate). The FieldSetMapper then turns that FieldSet into your domain object. Tokenizer = split; FieldSet = parsed fields; FieldSetMapper = bind to object.
- Your file has a bad row that fails parsing. How do you keep the job running?Make the step fault-tolerant and skip FlatFileParseException (e.g. .faultTolerant().skip(FlatFileParseException.class).skipLimit(n)), optionally logging via a SkipListener. Only the offending row is skipped, not the chunk.
saying these in an interview costs you the question
- Thinking FlatFileItemReader reads the whole file into memory at once (it streams line by line)
- Believing it is thread-safe by default
- Confusing the LineTokenizer (split) with the FieldSetMapper (bind)