skip to content

Flat-File Mapping & Tokenizing

Flat-file support tokenizes a line into fields and maps them onto an object, with aggregators and extractors doing the reverse on write. Interviewers use a fixed-width or CSV file as the setting for most Spring Batch exercises.

part ofSpring Frameworkoverview, primer and where to startread it →
on this pageshow

explore

questions

5

In Spring Batch, how does a FlatFileItemReader turn a line of text into a domain object? Name the pieces involved.

level: juniorimportance: must knowfreq 70%

answer

  1. raw line → tokenizer → FieldSet → mapper → object
  2. DefaultLineMapper = LineTokenizer + FieldSetMapper
  3. LineMapper.mapLine(line, lineNumber)
  4. FieldSet = typed, named/indexed tokens
  5. swap format = swap tokenizer only

basics

~20 s

The reader reads one raw line, then hands it to a LineMapper. The LineMapper uses a LineTokenizer to split the line into fields (a FieldSet), and a FieldSetMapper to turn those fields into a domain object.

solid answer

~30 s

FlatFileItemReader reads the file line by line as raw String text. Each line is passed to a LineMapper via mapLine(line, lineNumber). The default implementation, DefaultLineMapper, is a two-step pipeline: first a LineTokenizer (e.g. DelimitedLineTokenizer) splits the raw line into a FieldSet — an indexed, optionally named collection of typed string fields; then a FieldSetMapper (e.g. BeanWrapperFieldSetMapper) converts that FieldSet into a target object. So the flow is: raw line → LineTokenizer → FieldSet → FieldSetMapper → domain object. This separation lets you swap parsing (delimited vs fixed-length) independently from object binding.

code

java · 9 lines
java
FlatFileItemReader<Person> reader = new FlatFileItemReaderBuilder<Person>()
    .name("personReader")
    .resource(new ClassPathResource("people.csv"))
    .linesToSkip(1)                       // skip header
    .delimited()                          // DelimitedLineTokenizer
    .names("firstName", "lastName", "age")// FieldSet field names
    .targetType(Person.class)             // BeanWrapperFieldSetMapper
    .build();
// Under the hood: DefaultLineMapper(DelimitedLineTokenizer, BeanWrapperFieldSetMapper)

go deeper

for a junior

Must know the raw line → tokenizer → FieldSet → mapper → object chain and name LineMapper.

for a middle

Should name DefaultLineMapper and know how the builder wires tokenizer + mapper.

for a senior

Explains why decoupling tokenizing from mapping enables format/mapper reuse and how FlatFileParseException surfaces errors.

for a principal

Frames LineMapper as an extension seam for composite/heterogeneous formats and error-handling strategy.

## The read-side mapping pipeline Spring Batch's `FlatFileItemReader<T>` is an `ItemReader` that reads a resource (a file) one **line** at a time. Reading raw text and turning it into objects are two separate concerns, connected by the `LineMapper<T>` interface: ``` String line ──► LineMapper.mapLine(line, lineNumber) ──► T object ``` ### DefaultLineMapper — the standard two-stage pipeline `DefaultLineMapper<T>` is the usual implementation and composes two collaborators: 1. **`LineTokenizer`** — `FieldSet tokenize(String line)`. It splits one raw line into a **`FieldSet`**: an ordered collection of string tokens, each addressable by index (`readString(0)`) and, if names are configured, by name (`readString("firstName")`). `FieldSet` also offers typed accessors like `readInt`, `readBigDecimal`, `readDate`. Concrete tokenizers: `DelimitedLineTokenizer` (splits on a delimiter such as comma) and `FixedLengthTokenizer` (splits by column ranges). 2. **`FieldSetMapper<T>`** — `T mapFieldSet(FieldSet fieldSet)`. It binds the parsed fields onto a target object. The common implementation is `BeanWrapperFieldSetMapper`, which uses the field **names** to set JavaBean properties; you can also write a custom one for full control. ### Why the split matters Because tokenizing and mapping are decoupled, you can change the file format (delimited → fixed-width) by swapping only the tokenizer, keeping the same `FieldSetMapper`. Conversely you can reuse a tokenizer with different mappers. ### Terms defined - **FieldSet**: Spring Batch's typed wrapper around the array of string tokens from one record. It centralizes conversion (string → int/date/BigDecimal) so mappers don't parse strings manually. - **lineNumber**: the second argument to `mapLine`; useful for error messages. ### Edge cases / gotchas - Blank lines and header/comment lines are typically not tokenized: `FlatFileItemReader` has `setLinesToSkip(n)` (with an optional `SkippedLinesCallback`) and `setComments(new String[]{"#"})`. Tokenizing still fails loudly if a data line has the wrong number of columns (unless `setStrict(false)`). - A malformed line throws `FlatFileParseException` wrapping the underlying `LineTokenizer`/mapping error, carrying the line and line number. ### When to use Any CSV/TSV/fixed-width import. Use `FlatFileItemReaderBuilder` in modern config, which wires the `DefaultLineMapper`, tokenizer and mapper for you.

  • Which class implements LineMapper by default and what two objects does it delegate to?
    DefaultLineMapper — it delegates to a LineTokenizer (produces the FieldSet) and a FieldSetMapper (produces the object).
  • What is a FieldSet and why not just use a String[]?
    FieldSet wraps the tokens and adds name-based lookup and typed conversion (readInt, readDate, readBigDecimal), so mappers avoid manual string parsing and index bookkeeping.

saying these in an interview costs you the question

  • Thinking FlatFileItemReader parses straight into objects with no intermediate FieldSet
  • Confusing LineMapper (whole pipeline) with LineTokenizer (just the split step)
  • Believing the reader reads the whole file into memory at once

context

open as a page

Compare DelimitedLineTokenizer and FixedLengthTokenizer. How do you configure each, and what do field names give you?

level: middleimportance: must knowfreq 65%

basics

~10 s

DelimitedLineTokenizer splits a line on a delimiter like a comma. FixedLengthTokenizer splits by fixed column positions (Ranges). Both produce a FieldSet; setting names lets you read fields by name instead of index.

open as a page

On the write side, how does FlatFileItemWriter turn a domain object into a line? Explain LineAggregator and FieldExtractor.

level: seniorimportance: must knowfreq 55%

basics

~10 s

FlatFileItemWriter uses a LineAggregator to turn each object into a String line. A DelimitedLineAggregator (or FormatterLineAggregator) pulls values out via a FieldExtractor — usually BeanWrapperFieldExtractor with property names — then joins or formats them.

open as a page

How does BeanWrapperFieldSetMapper map a FieldSet to an object, and what are its requirements and limits?

level: middleimportance: should knowfreq 55%

basics

~10 s

BeanWrapperFieldSetMapper matches FieldSet field names to JavaBean properties and calls the setters, converting types automatically. The target class needs a no-arg constructor and matching setters. It can't populate constructor-only or immutable objects.

open as a page

You have a flat file with multiple record types (e.g. lines prefixed HEADER, TRADE, FOOTER) needing different tokenizers/mappers. How do you map it?

level: principalimportance: should knowfreq 35%

basics

~10 s

Use a PatternMatchingCompositeLineMapper. You register a tokenizer per line pattern (like HEADER*, TRADE*) and a FieldSetMapper per pattern. Each line is routed by its prefix to the right tokenizer and mapper.

open as a page