What is an ItemProcessor in Spring Batch, and where does it sit in a chunk-oriented step?
answer
- Middle stage: read -> process -> write
- One method O process(I)
- Optional; item-at-a-time
- Home for business logic
- Input type can differ from output type
basics
~10 sItemProcessor is the middle stage of a chunk step. For each item the ItemReader reads, its process() method transforms or filters it, then passes the result to the ItemWriter. It is optional.
solid answer
~40 sIn a chunk-oriented step, Spring Batch loops read -> process -> accumulate, then writes the whole chunk at once. ItemProcessor is the optional middle stage: for every item the ItemReader returns, the framework calls ItemProcessor.process(item), and whatever it returns is buffered and later handed to the ItemWriter. It has one method, O process(I item), where the input type I can differ from the output type O, so it commonly maps an input DTO to an output entity. It is the natural place for business logic: validation, enrichment, transformation, or filtering. If you configure a step with only reader and writer (no processor), items flow straight from reader to writer unchanged. It is called once per item, not once per chunk.
code
java · 17 linespublic class PersonProcessor implements ItemProcessor<PersonInput, PersonEntity> {
@Override
public PersonEntity process(PersonInput item) {
PersonEntity e = new PersonEntity();
e.setFullName(item.getFirstName() + " " + item.getLastName());
e.setEmail(item.getEmail().toLowerCase());
return e; // handed to the ItemWriter
}
}
// wiring
Step step = new StepBuilder("import", jobRepository)
.<PersonInput, PersonEntity>chunk(100, txManager)
.reader(reader)
.processor(new PersonProcessor())
.writer(writer)
.build();go deeper
Know it is the optional middle transform/filter stage, one method per item.
Explain item-at-a-time processing vs chunk-at-a-time writing, and I->O mapping.
Position it as the business-logic seam and mention filtering/exception semantics.
Discuss keeping processors stateless/idempotent and where cross-cutting concerns belong.
**Chunk-oriented processing.** A Spring Batch `Step` built with the chunk model works in a loop: the `ItemReader` is called repeatedly to read individual items until a chunk (e.g. `.chunk(100)`) is filled or the reader returns `null` (end of data). During that same loop each item is passed through the optional `ItemProcessor`, and the processed results are collected into a list. Once the chunk is complete, the whole list is passed **once** to the `ItemWriter.write(...)`, and the transaction commits. So processing is item-at-a-time, but writing is chunk-at-a-time. **The interface.** `org.springframework.batch.item.ItemProcessor<I, O>` has a single method: ```java @Nullable O process(@NonNull I item) throws Exception; ``` `I` is the input type (what the reader produces) and `O` is the output type (what the writer consumes). They can be the **same** type (transform in place) or **different** (map input model to output model). Because it is a functional interface, you can supply a lambda. **What it is for.** The processor is the designated home for business logic in a step: data transformation/mapping (e.g. CSV row -> JPA entity), enrichment (look up extra data from a service), validation, and filtering. Keeping this logic in the processor keeps the reader focused on I/O and the writer focused on persistence. **It is optional.** A step configured as reader + writer only is legal; items then pass through unchanged. Adding `.processor(myProcessor)` inserts the transform stage. **Ordering / call count.** `process()` is invoked exactly once per item read (before any filtering/skip logic drops it), in reading order, on the same thread as the reader unless you use multi-threaded or partitioned steps. **Return value semantics (preview).** Returning a non-null value passes it downstream; returning `null` **filters** the item out so it never reaches the writer (and increments the step's `filterCount`). Throwing an exception triggers skip/retry/rollback logic depending on step configuration. **When to use vs not.** Use a processor whenever the item must change shape, be enriched, or be conditionally dropped. Skip it when the reader already produces exactly what the writer needs.
- Is the ItemProcessor required in a chunk step?No. It is optional; with only a reader and writer, items flow through unchanged. You add it via .processor(...) when you need to transform, enrich, or filter.
- How many times is process() called per chunk?Once per item read into the chunk, in read order — not once for the whole chunk. The writer, by contrast, is called once per chunk with the whole list.
saying these in an interview costs you the question
- Saying the processor writes to the database (that's the writer's job)
- Claiming process() is called once per chunk
- Thinking the processor is mandatory