skip to content

Step Processing & Item Flow

How a step actually moves data: the chunk loop, readers, processors and writers, listeners, and the transaction boundary. This is where most practical Spring Batch questions live.

part ofSpring Frameworkoverview, primer and where to startread it →
on this pageshow

explore

questions

page 1 of 2

What is chunk-oriented processing in a Spring Batch step, and what does the commit-interval control?

level: juniorimportance: must knowfreq 75%

answer

  1. read-N, process-N, write-once, commit-once
  2. commit-interval = chunk size = n
  3. one transaction per chunk
  4. writer gets the whole list
  5. processor null = filter

basics

~20 s

A Spring Batch step reads items one by one, optionally processes each, and writes them in batches. The commit-interval (the chunk size) is how many items are grouped and written together in one transaction before committing.

solid answer

~40 s

Chunk-oriented processing is Spring Batch's default step model for high-volume data. The step reads items one at a time via an ItemReader, optionally transforms each via an ItemProcessor, and buffers them until it reaches the chunk size (the commit-interval). Then the whole buffered list is handed to the ItemWriter in a single write call, and the surrounding transaction commits. You configure it with StepBuilder.chunk(n, transactionManager). So chunk(10) means: read 10, process 10, write once (a list of 10), commit — then repeat. Writing in batches lets the writer do bulk operations (e.g. JDBC batch inserts) and keeps one transaction per chunk instead of one per row, which is what makes it scale.

code

java · 13 lines
java
@Bean
public Step importStep(JobRepository jobRepository,
                       PlatformTransactionManager transactionManager,
                       ItemReader<Person> reader,
                       ItemProcessor<Person, PersonDto> processor,
                       ItemWriter<PersonDto> writer) {
    return new StepBuilder("importStep", jobRepository)
            .<Person, PersonDto>chunk(100, transactionManager) // commit-interval = 100
            .reader(reader)       // read() called up to 100 times
            .processor(processor) // process(item) called per item
            .writer(writer)       // write(chunk) called ONCE per 100
            .build();
}

go deeper

for a junior

Must be able to state read-many/write-once and that commit-interval is the chunk size.

for a middle

Should mention one transaction per chunk and that the writer receives a list for bulk writes.

for a senior

Explains filtered-item counting, partial last chunk, and throughput reasoning behind batching the write.

for a principal

Frames chunk model as the scalability/restartability tradeoff vs per-row or whole-job transactions.

**The problem it solves.** Batch jobs move large volumes of data (millions of rows). Doing one transaction per row is catastrophically slow (commit overhead per item); doing one transaction for the whole job risks huge rollbacks, memory blowups, and no restartability. Spring Batch's *chunk-oriented processing* is the middle ground: process data in fixed-size groups called **chunks**, one transaction per chunk. **The three collaborators.** - **ItemReader<I>** — returns items one at a time via `read()`; returns `null` to signal end of input (e.g. `JdbcCursorItemReader`, `FlatFileItemReader`). - **ItemProcessor<I,O>** — optional; transforms/validates/filters one item at a time via `process(item)`. Returning `null` *filters* the item (it is skipped, not written). - **ItemWriter<O>** — writes a whole list at once via `write(Chunk<O> items)` (in Spring Batch 5; `List` in 4.x). Called **once per chunk**, enabling bulk/batch writes. **The loop.** For a commit-interval of `n`, one chunk iteration is: call `read()` up to `n` times (or until it returns `null`); for each item read, call `process(item)`; collect the non-null processed results; call `write(...)` once with that collected list; commit the transaction. Then start the next chunk. Concisely: **read-N, process-N, write-once, commit-once, repeat.** **commit-interval = chunk size = the `n`.** Historically the XML attribute was literally `commit-interval`. In the Java DSL you pass it to `chunk(n, transactionManager)`. It is the number of *items read* that trigger a commit — not the number written (filtered items still count toward the interval). **Why batch the write.** A single `write(list)` lets the `ItemWriter` issue a JDBC batch (`addBatch`/`executeBatch`), a single bulk HTTP call, etc. That amortizes per-round-trip cost across `n` items and is the main throughput lever. **Configuration (Spring Batch 5).** ```java new StepBuilder("step", jobRepository) .<Input, Output>chunk(100, transactionManager) .reader(reader).processor(processor).writer(writer).build(); ``` The `<Input, Output>` generics are the reader's input type and the writer's output type; the processor bridges them. **Edge cases / gotchas.** - **Last chunk is partial.** If the reader exhausts mid-chunk (read returns `null`), the accumulated items are written and committed even though fewer than `n`. - **Filtered items.** A processor returning `null` removes the item from the written list but it still counts toward the commit-interval read count. Filtered counts show up in step statistics as `filterCount`. - **Reader/processor are per-item; writer is per-chunk.** A common mistake is expecting the writer to be called once per item. - **Chunk size is a tuning knob.** Too small = commit overhead dominates; too large = big transactions, high memory, larger rollback/replay on failure. **When to use.** Chunk-oriented is the default and right choice for read→transform→write pipelines. The alternative, a **tasklet step** (`Tasklet.execute`), is for single non-item operations (run a stored proc, move a file, send a notification).

  • If chunk size is 100 but the file has 250 rows, how many times is the writer called and how big is each write?
    Three times: two writes of 100, then one final partial write of 50 when the reader returns null. Three transactions/commits total.
  • What happens to an item when the ItemProcessor returns null?
    It is filtered — excluded from the list passed to the writer — but it still counts toward the commit-interval. It is recorded as filterCount in the step execution statistics.

saying these in an interview costs you the question

  • Saying the writer is invoked once per item rather than once per chunk
  • Thinking filtered (processor-null) items don't count toward the commit-interval
  • Confusing commit-interval with a time interval

context

open as a page

In Spring Batch, how does a FlatFileItemReader turn a line of text into a domain object? Name the pieces involved.

level: juniorimportance: must knowfreq 70%

basics

~20 s

The reader reads one raw line, then hands it to a LineMapper. The LineMapper uses a LineTokenizer to split the line into fields (a FieldSet), and a FieldSetMapper to turn those fields into a domain object.

open as a page

What is an ItemProcessor in Spring Batch, and where does it sit in a chunk-oriented step?

level: juniorimportance: must knowfreq 70%

basics

~10 s

ItemProcessor is the middle stage of a chunk step. For each item the ItemReader reads, its process() method transforms or filters it, then passes the result to the ItemWriter. It is optional.

open as a page

What is the ItemReader interface in Spring Batch and what is the contract of its read() method?

level: juniorimportance: must knowfreq 75%

basics

~10 s

ItemReader is the input side of a chunk-oriented step. Its single method read() returns one item at a time, or null when there is no more data, signalling the end of input.

open as a page

What is an ItemWriter in Spring Batch, and why does its write method take a whole Chunk instead of a single item?

level: juniorimportance: must knowfreq 75%

basics

~20 s

An ItemWriter is the step component that outputs processed data. Its write method receives a Chunk (a batch of items) so it can write many items in one operation, like a single batched SQL insert, which is faster than one call per item.

open as a page

What is a Tasklet step in Spring Batch, and when would you use one instead of the chunk model?

level: juniorimportance: must knowfreq 70%

basics

~10 s

A Tasklet step runs a single custom action once, like a cleanup or a stored procedure call. You implement the Tasklet interface's execute() method. It's for tasks that aren't read-process-write loops over items.

open as a page

In a chunk-oriented Spring Batch step, what is the unit of work wrapped in a single transaction?

level: juniorimportance: must knowfreq 70%

basics

~10 s

One chunk. Spring Batch reads a group of items (the chunk size), processes them, then writes them all inside a single transaction. Commit happens once per chunk, not per item.

open as a page

Walk through the exact read/process/write sequence Spring Batch executes for one chunk, and where the transaction boundary sits.

level: middleimportance: must knowfreq 65%

basics

~20 s

A transaction starts, then the reader is called up to N times, the processor runs on each read item, the collected results are written in one write() call, and the transaction commits. Then the next chunk starts a new transaction.

open as a page

Compare DelimitedLineTokenizer and FixedLengthTokenizer. How do you configure each, and what do field names give you?

level: middleimportance: must knowfreq 65%

basics

~10 s

DelimitedLineTokenizer splits a line on a delimiter like a comma. FixedLengthTokenizer splits by fixed column positions (Ranges). Both produce a FieldSet; setting names lets you read fields by name instead of index.

open as a page

What happens when an ItemProcessor.process() returns null, and how is that different from throwing an exception?

level: middleimportance: must knowfreq 65%

basics

~10 s

Returning null filters the item: it is silently dropped and never reaches the writer, and Spring Batch increments the step's filterCount. Throwing an exception instead triggers skip/retry/rollback handling — a very different, error path.

open as a page

Explain RepeatStatus.FINISHED versus RepeatStatus.CONTINUABLE returned from Tasklet.execute().

level: middleimportance: must knowfreq 60%

basics

~10 s

FINISHED means the Tasklet is done and won't be called again. CONTINUABLE means there's more work, so Spring calls execute() again in a new transaction. Returning null counts as FINISHED.

open as a page

If writing the 50th item of a 100-item chunk throws, what happens to the other 99 items in that chunk?

level: middleimportance: must knowfreq 68%

basics

~10 s

The whole chunk's transaction rolls back, so none of the 100 items are committed. By default the step fails and stops. All 100 are undone together — the chunk is atomic.

open as a page

On the write side, how does FlatFileItemWriter turn a domain object into a line? Explain LineAggregator and FieldExtractor.

level: seniorimportance: must knowfreq 55%

basics

~10 s

FlatFileItemWriter uses a LineAggregator to turn each object into a String line. A DelimitedLineAggregator (or FormatterLineAggregator) pulls values out via a FieldExtractor — usually BeanWrapperFieldExtractor with property names — then joins or formats them.

open as a page

Explain the ItemStream interface (open/update/close) and how it enables restart of a failed Spring Batch step.

level: seniorimportance: must knowfreq 68%

basics

~20 s

ItemStream lets a reader/writer save and restore its position. open() initializes or restores state from the ExecutionContext, update() periodically writes the current position into it (at each chunk commit), and close() releases resources. On restart, open() reads the saved state so the step resumes where it stopped.

open as a page

How does BeanWrapperFieldSetMapper map a FieldSet to an object, and what are its requirements and limits?

level: middleimportance: should knowfreq 55%

basics

~10 s

BeanWrapperFieldSetMapper matches FieldSet field names to JavaBean properties and calls the setters, converting types automatically. The target class needs a no-arg constructor and matching setters. It can't populate constructor-only or immutable objects.

open as a page

How can an ItemProcessor change the item's type between reader and writer, and what constraints does that impose on step wiring?

level: middleimportance: should knowfreq 45%

basics

~20 s

ItemProcessor<I, O> can output a different type O than its input I. The reader must produce I, the processor maps I to O, and the writer must accept O. You declare both types on the step via chunk(...) generics.

open as a page

How do you configure a FlatFileItemReader to parse a delimited CSV file into domain objects?

level: middleimportance: should knowfreq 70%

basics

~20 s

FlatFileItemReader reads a text file line by line. You give it a Resource, a LineMapper that splits each line (a DelimitedLineTokenizer into a FieldSet) and maps the fields to an object (a FieldSetMapper), plus optional linesToSkip for headers.

open as a page

How do you configure a FlatFileItemWriter to produce a CSV, and what are the key options (resource, line aggregator, header/footer, append/restart)?

level: middleimportance: should knowfreq 55%

basics

~20 s

FlatFileItemWriter writes items as text lines to a file Resource. You give it a LineAggregator (e.g. DelimitedLineAggregator with a FieldExtractor) to turn each item into a line, and optionally a header/footer callback. It buffers lines and flushes per chunk.

open as a page

How can you make chunk size dynamic instead of fixed, and what role does CompletionPolicy play?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Pass a CompletionPolicy to chunk() instead of a fixed number. The policy decides after each read whether the chunk is complete — for example by item count, elapsed time, or custom logic — instead of always committing at a fixed N.

open as a page

What is CompositeItemProcessor and how does chaining behave with respect to types, ordering, and null-filtering?

level: seniorimportance: should knowfreq 40%

basics

~20 s

CompositeItemProcessor is an ItemProcessor that holds an ordered list of delegate processors and runs them in sequence, feeding each one's output into the next. If any delegate returns null, the chain stops and the item is filtered.

open as a page

When reading from a database in Spring Batch, when would you choose a cursor-based reader (JdbcCursorItemReader) over a paging reader (JpaPagingItemReader), and what are the trade-offs?

level: seniorimportance: should knowfreq 65%

basics

~20 s

Cursor readers stream one big result set over a single open connection/ResultSet — low memory, one query, but the connection stays open and it is not restartable across threads. Paging readers run repeated LIMIT/OFFSET queries per page — connection released between pages, restartable and partition-friendly, but more queries and needs a stable sort.

open as a page

When would you use CompositeItemWriter versus ClassifierCompositeItemWriter, and how do their write semantics differ?

level: seniorimportance: should knowfreq 40%

basics

~20 s

CompositeItemWriter sends every item to all its delegate writers in order (fan-out — write to DB and a file). ClassifierCompositeItemWriter uses a Classifier to route each item to one writer based on the item, so different items go to different destinations.

open as a page

Compare JdbcBatchItemWriter and JpaItemWriter: how each persists a chunk, and when you would choose one over the other.

level: seniorimportance: should knowfreq 50%

basics

~20 s

JdbcBatchItemWriter runs a single parameterized SQL statement as a JDBC batch for the whole chunk. JpaItemWriter persists or merges each entity through an EntityManager and flushes once. Choose JDBC for raw speed on plain SQL, JPA when you already work with managed entities.

open as a page

What is MethodInvokingTaskletAdapter and when would you use it?

level: seniorimportance: should knowfreq 35%

basics

~20 s

It's an adapter that turns any existing Spring bean method into a Tasklet, so you don't have to implement the Tasklet interface. You set the target object and the method name, and its return value can drive exit status.

open as a page

How does SystemCommandTasklet work and what must you configure to use it safely?

level: seniorimportance: should knowfreq 30%

basics

~20 s

SystemCommandTasklet runs an external OS command (like a shell script) as a batch step. You set the command, a timeout, and usually a working directory. It runs the command in a separate thread and checks the exit code.

open as a page

You have a Spring Batch step that reads a CSV and calls a REST API with no database at all. Which PlatformTransactionManager should the chunk use, and why?

level: seniorimportance: should knowfreq 40%

basics

~10 s

Use ResourcelessTransactionManager. Chunk steps require a PlatformTransactionManager, but with no database there is no real transactional resource to manage, so this no-op manager satisfies the API without touching a datasource.

open as a page

What is the transactionAttribute on a chunk step, and what can you tune with it?

level: seniorimportance: should knowfreq 45%

basics

~10 s

It configures the chunk transaction's properties — propagation, isolation level, timeout, and read-only flag — via a DefaultTransactionAttribute. You set it on the StepBuilder to override defaults like isolation or a per-chunk timeout.

open as a page

How do you choose a chunk (commit-interval) size, and what are the tradeoffs of too-large vs too-small values?

level: principalimportance: should knowfreq 45%

basics

~20 s

Bigger chunks mean fewer commits and faster throughput but more memory, longer transactions, and more work re-done on failure. Smaller chunks mean more commit overhead but lower memory and cheaper retries. You tune empirically for your data and DB.

open as a page

You have a flat file with multiple record types (e.g. lines prefixed HEADER, TRADE, FOOTER) needing different tokenizers/mappers. How do you map it?

level: principalimportance: should knowfreq 35%

basics

~10 s

Use a PatternMatchingCompositeLineMapper. You register a tokenizer per line pattern (like HEADER*, TRADE*) and a FieldSetMapper per pattern. Each line is routed by its prefix to the right tokenizer and mapper.

open as a page

What design constraints (statelessness, idempotency, side effects) apply to an ItemProcessor in a fault-tolerant or multi-threaded step, and why?

level: principalimportance: should knowfreq 30%

basics

~20 s

Keep processors stateless, side-effect-free, and idempotent. In fault-tolerant steps a chunk rollback causes items to be re-processed, so process() may be called more than once per item; in multi-threaded steps the same instance runs concurrently, so mutable fields are unsafe.

open as a page

showing 1–30 of 34