You have a flat file with multiple record types (e.g. lines prefixed HEADER, TRADE, FOOTER) needing different tokenizers/mappers. How do you map it?
answer
- PatternMatchingCompositeLineMapper = pattern→tokenizer + pattern→mapper
- keys like HEADER*, TRADE*, * (PatternMatcher wildcards)
- common supertype for T; concrete subtype per mapper
- write side: ClassifierCompositeItemWriter
- add "*" fallback; maps must share keys
basics
~10 sUse a PatternMatchingCompositeLineMapper. You register a tokenizer per line pattern (like HEADER*, TRADE*) and a FieldSetMapper per pattern. Each line is routed by its prefix to the right tokenizer and mapper.
solid answer
~40 sFor heterogeneous files you can't use a single DelimitedLineTokenizer. Spring Batch provides PatternMatchingCompositeLineMapper, a LineMapper that holds a Map<String, LineTokenizer> and a Map<String, FieldSetMapper> keyed by line patterns (using PatternMatcher wildcards like "HEADER*", "TRADE*", "*"). For each line it matches the pattern, tokenizes with the matching tokenizer, then maps the FieldSet with the matching mapper — so different record types get different parsing and produce different object types (usually a common supertype). Downstream, your processor/writer inspects the object type (or you use a ClassifierCompositeItemWriter to route each type to the right writer). Alternatively, write a fully custom LineMapper. This keeps mixed-format files (header/detail/trailer, multi-line records) manageable while reusing the standard tokenizer/mapper components.
code
java · 16 linesPatternMatchingCompositeLineMapper<Record> mapper = new PatternMatchingCompositeLineMapper<>();
Map<String, LineTokenizer> tokenizers = new HashMap<>();
tokenizers.put("HEADER*", headerTokenizer);
tokenizers.put("TRADE*", tradeTokenizer);
tokenizers.put("FOOTER*", footerTokenizer);
mapper.setTokenizers(tokenizers);
Map<String, FieldSetMapper<Record>> mappers = new HashMap<>();
mappers.put("HEADER*", headerMapper);
mappers.put("TRADE*", tradeMapper);
mappers.put("FOOTER*", footerMapper);
mapper.setFieldSetMappers(mappers);
// Reader uses this LineMapper; each line routed by prefix
reader.setLineMapper(mapper);go deeper
Likely unaware; may only know a single tokenizer.
Knows the problem exists and that a composite/custom LineMapper is needed.
Configures PatternMatchingCompositeLineMapper with pattern-keyed maps and a shared supertype.
Designs the full in/out pipeline (composite mapper + ClassifierCompositeItemWriter/Processor), handles unmatched lines and multi-line grouping trade-offs.
## Mapping files with more than one record shape Real-world flat files are often **heterogeneous**: a `HEADER` line, many `TRADE`/detail lines with a different layout, and a `FOOTER`/trailer with counts. One `DelimitedLineTokenizer` can't handle all three because they have different columns and meanings. ### PatternMatchingCompositeLineMapper `PatternMatchingCompositeLineMapper<T>` implements `LineMapper<T>` and routes each line by a **prefix/pattern**: - `setTokenizers(Map<String, LineTokenizer>)` — pattern → tokenizer. Keys use Spring Batch's `PatternMatcher` wildcard syntax (`*`, `?`), e.g. `"HEADER*"`, `"TRADE*"`, and a catch-all `"*"`. - `setFieldSetMappers(Map<String, FieldSetMapper>)` — pattern → mapper, same keys. For each raw line it: (1) finds the tokenizer whose pattern matches the line, (2) tokenizes to a `FieldSet`, (3) finds the matching `FieldSetMapper`, (4) maps to an object. Typically all record types share a **common supertype/interface** so the reader's generic type `T` works, and each mapper produces the concrete subtype. ``` line "TRADE|AAPL|100" ─match "TRADE*"─► tradeTokenizer ─► FieldSet ─► tradeMapper ─► Trade line "HEADER|2026..." ─match "HEADER*"─► headerTokenizer ─► FieldSet ─► headerMapper ─► Header ``` ### Routing on the write side When records are heterogeneous going **out**, mirror with a `ClassifierCompositeItemWriter` (a `Classifier` maps each item type to the appropriate delegate `FlatFileItemWriter`), and a `ClassifierCompositeItemProcessor` if processing differs per type. ### Multi-line records When one logical record spans several physical lines (e.g. a header line followed by N detail lines that belong together), `PatternMatchingCompositeLineMapper` alone isn't enough — wrap a delegate reader in an `AggregateItemReader`-style/custom reader or a `MultiResourceItemReader` pattern that groups lines. That crosses into reader lifecycle (owned by a sibling leaf), but the **mapping** of each individual line still uses the composite mapper. ### Alternatives & trade-offs - **Custom `LineMapper`**: implement `mapLine(line, lineNumber)` yourself and branch on the prefix. Maximum flexibility; more code; you lose the declarative pattern map. - **PatternMatchingCompositeLineMapper**: declarative, reuses standard tokenizers/mappers, self-documenting patterns. Preferred for straightforward prefix-routed formats. ### Gotchas - Patterns must be **exhaustive** — an unmatched line throws; add a `"*"` fallback if you want to skip/handle unknowns. - The tokenizer map and mapper map must use the **same keys**. - All produced types must be assignable to the reader's `T` (use a shared marker interface). - Don't forget that skipping/comment handling still happens at the `FlatFileItemReader` level before mapping. ### When to use Header/detail/trailer files, mixed transaction feeds, EDI-like line-typed formats — any single file where the layout depends on a line's type indicator.
- What generic type would the reader use if HEADER, TRADE and FOOTER map to different classes?A common supertype or marker interface (e.g. Record) that all three concrete classes implement, so FlatFileItemReader<Record> and the per-pattern mappers agree on T.
- How do you write these heterogeneous objects back out to (possibly different) files?Use a ClassifierCompositeItemWriter with a Classifier that routes each item type to a dedicated delegate FlatFileItemWriter (each with its own LineAggregator).
- What happens if a line matches none of the configured patterns?It throws (no tokenizer/mapper found). Add a catch-all "*" pattern to handle or skip unknown lines gracefully.
saying these in an interview costs you the question
- Trying to force one DelimitedLineTokenizer to parse all record types
- Forgetting the tokenizer and mapper maps must share the same pattern keys
- Not providing a common supertype for the reader's generic parameter
- Assuming the composite mapper also groups multi-line logical records (that's reader-level)