In Spring Batch, how does a FlatFileItemReader turn a line of text into a domain object? Name the pieces involved.
answer
- raw line → tokenizer → FieldSet → mapper → object
- DefaultLineMapper = LineTokenizer + FieldSetMapper
- LineMapper.mapLine(line, lineNumber)
- FieldSet = typed, named/indexed tokens
- swap format = swap tokenizer only
basics
~20 sThe reader reads one raw line, then hands it to a LineMapper. The LineMapper uses a LineTokenizer to split the line into fields (a FieldSet), and a FieldSetMapper to turn those fields into a domain object.
solid answer
~30 sFlatFileItemReader reads the file line by line as raw String text. Each line is passed to a LineMapper via mapLine(line, lineNumber). The default implementation, DefaultLineMapper, is a two-step pipeline: first a LineTokenizer (e.g. DelimitedLineTokenizer) splits the raw line into a FieldSet — an indexed, optionally named collection of typed string fields; then a FieldSetMapper (e.g. BeanWrapperFieldSetMapper) converts that FieldSet into a target object. So the flow is: raw line → LineTokenizer → FieldSet → FieldSetMapper → domain object. This separation lets you swap parsing (delimited vs fixed-length) independently from object binding.
code
java · 9 linesFlatFileItemReader<Person> reader = new FlatFileItemReaderBuilder<Person>()
.name("personReader")
.resource(new ClassPathResource("people.csv"))
.linesToSkip(1) // skip header
.delimited() // DelimitedLineTokenizer
.names("firstName", "lastName", "age")// FieldSet field names
.targetType(Person.class) // BeanWrapperFieldSetMapper
.build();
// Under the hood: DefaultLineMapper(DelimitedLineTokenizer, BeanWrapperFieldSetMapper)go deeper
Must know the raw line → tokenizer → FieldSet → mapper → object chain and name LineMapper.
Should name DefaultLineMapper and know how the builder wires tokenizer + mapper.
Explains why decoupling tokenizing from mapping enables format/mapper reuse and how FlatFileParseException surfaces errors.
Frames LineMapper as an extension seam for composite/heterogeneous formats and error-handling strategy.
## The read-side mapping pipeline Spring Batch's `FlatFileItemReader<T>` is an `ItemReader` that reads a resource (a file) one **line** at a time. Reading raw text and turning it into objects are two separate concerns, connected by the `LineMapper<T>` interface: ``` String line ──► LineMapper.mapLine(line, lineNumber) ──► T object ``` ### DefaultLineMapper — the standard two-stage pipeline `DefaultLineMapper<T>` is the usual implementation and composes two collaborators: 1. **`LineTokenizer`** — `FieldSet tokenize(String line)`. It splits one raw line into a **`FieldSet`**: an ordered collection of string tokens, each addressable by index (`readString(0)`) and, if names are configured, by name (`readString("firstName")`). `FieldSet` also offers typed accessors like `readInt`, `readBigDecimal`, `readDate`. Concrete tokenizers: `DelimitedLineTokenizer` (splits on a delimiter such as comma) and `FixedLengthTokenizer` (splits by column ranges). 2. **`FieldSetMapper<T>`** — `T mapFieldSet(FieldSet fieldSet)`. It binds the parsed fields onto a target object. The common implementation is `BeanWrapperFieldSetMapper`, which uses the field **names** to set JavaBean properties; you can also write a custom one for full control. ### Why the split matters Because tokenizing and mapping are decoupled, you can change the file format (delimited → fixed-width) by swapping only the tokenizer, keeping the same `FieldSetMapper`. Conversely you can reuse a tokenizer with different mappers. ### Terms defined - **FieldSet**: Spring Batch's typed wrapper around the array of string tokens from one record. It centralizes conversion (string → int/date/BigDecimal) so mappers don't parse strings manually. - **lineNumber**: the second argument to `mapLine`; useful for error messages. ### Edge cases / gotchas - Blank lines and header/comment lines are typically not tokenized: `FlatFileItemReader` has `setLinesToSkip(n)` (with an optional `SkippedLinesCallback`) and `setComments(new String[]{"#"})`. Tokenizing still fails loudly if a data line has the wrong number of columns (unless `setStrict(false)`). - A malformed line throws `FlatFileParseException` wrapping the underlying `LineTokenizer`/mapping error, carrying the line and line number. ### When to use Any CSV/TSV/fixed-width import. Use `FlatFileItemReaderBuilder` in modern config, which wires the `DefaultLineMapper`, tokenizer and mapper for you.
- Which class implements LineMapper by default and what two objects does it delegate to?DefaultLineMapper — it delegates to a LineTokenizer (produces the FieldSet) and a FieldSetMapper (produces the object).
- What is a FieldSet and why not just use a String[]?FieldSet wraps the tokens and adds name-based lookup and typed conversion (readInt, readDate, readBigDecimal), so mappers avoid manual string parsing and index bookkeeping.
saying these in an interview costs you the question
- Thinking FlatFileItemReader parses straight into objects with no intermediate FieldSet
- Confusing LineMapper (whole pipeline) with LineTokenizer (just the split step)
- Believing the reader reads the whole file into memory at once