skip to content

ItemProcessor

The processor transforms an item, may change its type, and drops it entirely by returning null. Interviewers ask what returning null does, since silently filtered rows are a classic batch mystery.

part ofSpring Frameworkoverview, primer and where to startread it →
on this pageshow

explore

questions

5

What is an ItemProcessor in Spring Batch, and where does it sit in a chunk-oriented step?

level: juniorimportance: must knowfreq 70%

answer

  1. Middle stage: read -> process -> write
  2. One method O process(I)
  3. Optional; item-at-a-time
  4. Home for business logic
  5. Input type can differ from output type

basics

~10 s

ItemProcessor is the middle stage of a chunk step. For each item the ItemReader reads, its process() method transforms or filters it, then passes the result to the ItemWriter. It is optional.

solid answer

~40 s

In a chunk-oriented step, Spring Batch loops read -> process -> accumulate, then writes the whole chunk at once. ItemProcessor is the optional middle stage: for every item the ItemReader returns, the framework calls ItemProcessor.process(item), and whatever it returns is buffered and later handed to the ItemWriter. It has one method, O process(I item), where the input type I can differ from the output type O, so it commonly maps an input DTO to an output entity. It is the natural place for business logic: validation, enrichment, transformation, or filtering. If you configure a step with only reader and writer (no processor), items flow straight from reader to writer unchanged. It is called once per item, not once per chunk.

code

java · 17 lines
java
public class PersonProcessor implements ItemProcessor<PersonInput, PersonEntity> {
    @Override
    public PersonEntity process(PersonInput item) {
        PersonEntity e = new PersonEntity();
        e.setFullName(item.getFirstName() + " " + item.getLastName());
        e.setEmail(item.getEmail().toLowerCase());
        return e; // handed to the ItemWriter
    }
}

// wiring
Step step = new StepBuilder("import", jobRepository)
        .<PersonInput, PersonEntity>chunk(100, txManager)
        .reader(reader)
        .processor(new PersonProcessor())
        .writer(writer)
        .build();

go deeper

for a junior

Know it is the optional middle transform/filter stage, one method per item.

for a middle

Explain item-at-a-time processing vs chunk-at-a-time writing, and I->O mapping.

for a senior

Position it as the business-logic seam and mention filtering/exception semantics.

for a principal

Discuss keeping processors stateless/idempotent and where cross-cutting concerns belong.

**Chunk-oriented processing.** A Spring Batch `Step` built with the chunk model works in a loop: the `ItemReader` is called repeatedly to read individual items until a chunk (e.g. `.chunk(100)`) is filled or the reader returns `null` (end of data). During that same loop each item is passed through the optional `ItemProcessor`, and the processed results are collected into a list. Once the chunk is complete, the whole list is passed **once** to the `ItemWriter.write(...)`, and the transaction commits. So processing is item-at-a-time, but writing is chunk-at-a-time. **The interface.** `org.springframework.batch.item.ItemProcessor<I, O>` has a single method: ```java @Nullable O process(@NonNull I item) throws Exception; ``` `I` is the input type (what the reader produces) and `O` is the output type (what the writer consumes). They can be the **same** type (transform in place) or **different** (map input model to output model). Because it is a functional interface, you can supply a lambda. **What it is for.** The processor is the designated home for business logic in a step: data transformation/mapping (e.g. CSV row -> JPA entity), enrichment (look up extra data from a service), validation, and filtering. Keeping this logic in the processor keeps the reader focused on I/O and the writer focused on persistence. **It is optional.** A step configured as reader + writer only is legal; items then pass through unchanged. Adding `.processor(myProcessor)` inserts the transform stage. **Ordering / call count.** `process()` is invoked exactly once per item read (before any filtering/skip logic drops it), in reading order, on the same thread as the reader unless you use multi-threaded or partitioned steps. **Return value semantics (preview).** Returning a non-null value passes it downstream; returning `null` **filters** the item out so it never reaches the writer (and increments the step's `filterCount`). Throwing an exception triggers skip/retry/rollback logic depending on step configuration. **When to use vs not.** Use a processor whenever the item must change shape, be enriched, or be conditionally dropped. Skip it when the reader already produces exactly what the writer needs.

  • Is the ItemProcessor required in a chunk step?
    No. It is optional; with only a reader and writer, items flow through unchanged. You add it via .processor(...) when you need to transform, enrich, or filter.
  • How many times is process() called per chunk?
    Once per item read into the chunk, in read order — not once for the whole chunk. The writer, by contrast, is called once per chunk with the whole list.

saying these in an interview costs you the question

  • Saying the processor writes to the database (that's the writer's job)
  • Claiming process() is called once per chunk
  • Thinking the processor is mandatory

context

open as a page

What happens when an ItemProcessor.process() returns null, and how is that different from throwing an exception?

level: middleimportance: must knowfreq 65%

basics

~10 s

Returning null filters the item: it is silently dropped and never reaches the writer, and Spring Batch increments the step's filterCount. Throwing an exception instead triggers skip/retry/rollback handling — a very different, error path.

open as a page

How can an ItemProcessor change the item's type between reader and writer, and what constraints does that impose on step wiring?

level: middleimportance: should knowfreq 45%

basics

~20 s

ItemProcessor<I, O> can output a different type O than its input I. The reader must produce I, the processor maps I to O, and the writer must accept O. You declare both types on the step via chunk(...) generics.

open as a page

What is CompositeItemProcessor and how does chaining behave with respect to types, ordering, and null-filtering?

level: seniorimportance: should knowfreq 40%

basics

~20 s

CompositeItemProcessor is an ItemProcessor that holds an ordered list of delegate processors and runs them in sequence, feeding each one's output into the next. If any delegate returns null, the chain stops and the item is filtered.

open as a page

What design constraints (statelessness, idempotency, side effects) apply to an ItemProcessor in a fault-tolerant or multi-threaded step, and why?

level: principalimportance: should knowfreq 30%

basics

~20 s

Keep processors stateless, side-effect-free, and idempotent. In fault-tolerant steps a chunk rollback causes items to be re-processed, so process() may be called more than once per item; in multi-threaded steps the same instance runs concurrently, so mutable fields are unsafe.

open as a page