skip to content

You have a flat file with multiple record types (e.g. lines prefixed HEADER, TRADE, FOOTER) needing different tokenizers/mappers. How do you map it?

level: principalimportance: should knowfreq 35%

answer

  1. PatternMatchingCompositeLineMapper = pattern→tokenizer + pattern→mapper
  2. keys like HEADER*, TRADE*, * (PatternMatcher wildcards)
  3. common supertype for T; concrete subtype per mapper
  4. write side: ClassifierCompositeItemWriter
  5. add "*" fallback; maps must share keys

basics

~10 s

Use a PatternMatchingCompositeLineMapper. You register a tokenizer per line pattern (like HEADER*, TRADE*) and a FieldSetMapper per pattern. Each line is routed by its prefix to the right tokenizer and mapper.

solid answer

~40 s

For heterogeneous files you can't use a single DelimitedLineTokenizer. Spring Batch provides PatternMatchingCompositeLineMapper, a LineMapper that holds a Map<String, LineTokenizer> and a Map<String, FieldSetMapper> keyed by line patterns (using PatternMatcher wildcards like "HEADER*", "TRADE*", "*"). For each line it matches the pattern, tokenizes with the matching tokenizer, then maps the FieldSet with the matching mapper — so different record types get different parsing and produce different object types (usually a common supertype). Downstream, your processor/writer inspects the object type (or you use a ClassifierCompositeItemWriter to route each type to the right writer). Alternatively, write a fully custom LineMapper. This keeps mixed-format files (header/detail/trailer, multi-line records) manageable while reusing the standard tokenizer/mapper components.

code

java · 16 lines
java
PatternMatchingCompositeLineMapper<Record> mapper = new PatternMatchingCompositeLineMapper<>();

Map<String, LineTokenizer> tokenizers = new HashMap<>();
tokenizers.put("HEADER*", headerTokenizer);
tokenizers.put("TRADE*",  tradeTokenizer);
tokenizers.put("FOOTER*", footerTokenizer);
mapper.setTokenizers(tokenizers);

Map<String, FieldSetMapper<Record>> mappers = new HashMap<>();
mappers.put("HEADER*", headerMapper);
mappers.put("TRADE*",  tradeMapper);
mappers.put("FOOTER*", footerMapper);
mapper.setFieldSetMappers(mappers);

// Reader uses this LineMapper; each line routed by prefix
reader.setLineMapper(mapper);

go deeper

for a junior

Likely unaware; may only know a single tokenizer.

for a middle

Knows the problem exists and that a composite/custom LineMapper is needed.

for a senior

Configures PatternMatchingCompositeLineMapper with pattern-keyed maps and a shared supertype.

for a principal

Designs the full in/out pipeline (composite mapper + ClassifierCompositeItemWriter/Processor), handles unmatched lines and multi-line grouping trade-offs.

## Mapping files with more than one record shape Real-world flat files are often **heterogeneous**: a `HEADER` line, many `TRADE`/detail lines with a different layout, and a `FOOTER`/trailer with counts. One `DelimitedLineTokenizer` can't handle all three because they have different columns and meanings. ### PatternMatchingCompositeLineMapper `PatternMatchingCompositeLineMapper<T>` implements `LineMapper<T>` and routes each line by a **prefix/pattern**: - `setTokenizers(Map<String, LineTokenizer>)` — pattern → tokenizer. Keys use Spring Batch's `PatternMatcher` wildcard syntax (`*`, `?`), e.g. `"HEADER*"`, `"TRADE*"`, and a catch-all `"*"`. - `setFieldSetMappers(Map<String, FieldSetMapper>)` — pattern → mapper, same keys. For each raw line it: (1) finds the tokenizer whose pattern matches the line, (2) tokenizes to a `FieldSet`, (3) finds the matching `FieldSetMapper`, (4) maps to an object. Typically all record types share a **common supertype/interface** so the reader's generic type `T` works, and each mapper produces the concrete subtype. ``` line "TRADE|AAPL|100" ─match "TRADE*"─► tradeTokenizer ─► FieldSet ─► tradeMapper ─► Trade line "HEADER|2026..." ─match "HEADER*"─► headerTokenizer ─► FieldSet ─► headerMapper ─► Header ``` ### Routing on the write side When records are heterogeneous going **out**, mirror with a `ClassifierCompositeItemWriter` (a `Classifier` maps each item type to the appropriate delegate `FlatFileItemWriter`), and a `ClassifierCompositeItemProcessor` if processing differs per type. ### Multi-line records When one logical record spans several physical lines (e.g. a header line followed by N detail lines that belong together), `PatternMatchingCompositeLineMapper` alone isn't enough — wrap a delegate reader in an `AggregateItemReader`-style/custom reader or a `MultiResourceItemReader` pattern that groups lines. That crosses into reader lifecycle (owned by a sibling leaf), but the **mapping** of each individual line still uses the composite mapper. ### Alternatives & trade-offs - **Custom `LineMapper`**: implement `mapLine(line, lineNumber)` yourself and branch on the prefix. Maximum flexibility; more code; you lose the declarative pattern map. - **PatternMatchingCompositeLineMapper**: declarative, reuses standard tokenizers/mappers, self-documenting patterns. Preferred for straightforward prefix-routed formats. ### Gotchas - Patterns must be **exhaustive** — an unmatched line throws; add a `"*"` fallback if you want to skip/handle unknowns. - The tokenizer map and mapper map must use the **same keys**. - All produced types must be assignable to the reader's `T` (use a shared marker interface). - Don't forget that skipping/comment handling still happens at the `FlatFileItemReader` level before mapping. ### When to use Header/detail/trailer files, mixed transaction feeds, EDI-like line-typed formats — any single file where the layout depends on a line's type indicator.

  • What generic type would the reader use if HEADER, TRADE and FOOTER map to different classes?
    A common supertype or marker interface (e.g. Record) that all three concrete classes implement, so FlatFileItemReader<Record> and the per-pattern mappers agree on T.
  • How do you write these heterogeneous objects back out to (possibly different) files?
    Use a ClassifierCompositeItemWriter with a Classifier that routes each item type to a dedicated delegate FlatFileItemWriter (each with its own LineAggregator).
  • What happens if a line matches none of the configured patterns?
    It throws (no tokenizer/mapper found). Add a catch-all "*" pattern to handle or skip unknown lines gracefully.

saying these in an interview costs you the question

  • Trying to force one DelimitedLineTokenizer to parse all record types
  • Forgetting the tokenizer and mapper maps must share the same pattern keys
  • Not providing a common supertype for the reader's generic parameter
  • Assuming the composite mapper also groups multi-line logical records (that's reader-level)

context