How can an ItemProcessor change the item's type between reader and writer, and what constraints does that impose on step wiring?
answer
- ItemProcessor<I, O> — two type params
- Read DTO -> write entity
- .<I,O>chunk(...) generics
- Mismatch = ClassCastException at runtime
- Composite chains intermediate types
basics
~20 sItemProcessor<I, O> can output a different type O than its input I. The reader must produce I, the processor maps I to O, and the writer must accept O. You declare both types on the step via chunk(...) generics.
solid answer
~40 sThe interface is ItemProcessor<I, O>, so the output type O is independent of the input type I. This is the standard read-model to write-model mapping: the ItemReader emits type I (e.g. a flat-file DTO), the processor converts each to type O (e.g. a JPA entity), and the ItemWriter must be typed for O. The key wiring constraint is that the step is parameterized as .<I, O>chunk(size, txManager): the first generic is the reader's output/processor's input, the second is the processor's output/writer's input. Mismatches surface as compile-time generic errors or ClassCastExceptions at runtime. If you have no processor, I and O collapse to the same type. When you need several transformation stages producing intermediate types, chain them with CompositeItemProcessor, where each stage's output feeds the next stage's input.
code
java · 19 lines// Reader yields TradeCsv, writer persists TradeEntity — processor bridges the types
public class TradeMapper implements ItemProcessor<TradeCsv, TradeEntity> {
@Override
public TradeEntity process(TradeCsv row) {
return new TradeEntity(
row.getIsin(),
new BigDecimal(row.getAmount()),
LocalDate.parse(row.getTradeDate()));
}
}
@Bean
Step tradeStep(JobRepository repo, PlatformTransactionManager tx,
ItemReader<TradeCsv> reader, ItemWriter<TradeEntity> writer) {
return new StepBuilder("tradeStep", repo)
.<TradeCsv, TradeEntity>chunk(100, tx)
.reader(reader).processor(new TradeMapper()).writer(writer)
.build();
}go deeper
Know the output type can differ from the input type.
Explain the <I,O> chunk generics and reader/writer type alignment.
Discuss composite chains of intermediate types and erasure-driven runtime failures.
Weigh read-model/write-model separation and idempotent pure mapping under retry.
**Two type parameters.** `ItemProcessor<I, O>` deliberately separates input and output types. `I` is whatever the `ItemReader<I>` produces; `O` is whatever the `ItemWriter<O>` (in Spring Batch 5, `ItemWriter<? super O>`) consumes. The processor is the type-conversion seam of the step. **Why change type.** The most common batch shape is *read one model, write another*: read a CSV/flat-file record into a lightweight input DTO, then map it to a persistence entity or an outbound message DTO. Keeping the read model and write model separate keeps parsing concerns out of your domain entity and lets validation/enrichment happen in between. **Step wiring.** The chunk builder carries both generics: ```java new StepBuilder("step", jobRepository) .<InputDto, OutputEntity>chunk(100, txManager) // <I, O> .reader(inputDtoReader) // ItemReader<InputDto> .processor(mapper) // ItemProcessor<InputDto, OutputEntity> .writer(entityWriter) // ItemWriter<OutputEntity> .build(); ``` The first type argument (`InputDto`) must match the reader's item type and the processor's input; the second (`OutputEntity`) must match the processor's output and the writer's input. Get these wrong and you get either a **compile error** (with typed beans) or a **ClassCastException** at runtime (with raw/loosely typed beans or `PassThroughItemProcessor` misuse). **No processor case.** If the step has no processor, there is effectively one type: reader output == writer input, and you often write `.<T, T>chunk(...)`. **Multiple conversions.** When conversion needs several stages of differing types, use `CompositeItemProcessor`. Its own generics are `<I, O>` for the overall chain, but internally each delegate's output type must be assignable to the next delegate's input type. The compiler cannot fully verify the internal chain because the delegate list is type-erased, so a mismatch there manifests as a runtime `ClassCastException` — a well-known gotcha. **Edge cases / gotchas.** - Because `O` can differ from `I`, a filtered item (return null) simply means 'no `O` produced for this `I`'. - Generics on Spring Batch beans are erased at runtime; the framework does not enforce type safety beyond what the compiler sees, so integration tests matter. - Under a fault-tolerant retry the processor may re-run on the same `I` item, so the `I -> O` mapping should be pure/idempotent (no side effects, no mutating the input in a way that breaks a re-run). **When to use.** Any time the persisted/written shape differs from the read shape — which is most real ETL jobs.
- Where do you declare the input and output types for a type-changing step?On the chunk builder: .<I, O>chunk(size, txManager). The first generic is reader-output/processor-input, the second is processor-output/writer-input.
- Why might a CompositeItemProcessor with mismatched delegate types compile fine but fail at runtime?The delegate list is type-erased, so the compiler can't verify that each stage's output matches the next stage's input; a mismatch surfaces as a ClassCastException when the chain runs.
saying these in an interview costs you the question
- Claiming input and output types must be identical
- Not knowing where the <I, O> generics are declared
- Thinking the framework enforces delegate chain types at runtime