skip to content

Explain the ItemStream interface (open/update/close) and how it enables restart of a failed Spring Batch step.

level: seniorimportance: must knowfreq 68%

answer

  1. open = init or restore, update = checkpoint per chunk, close = release
  2. state lives in ExecutionContext → JobRepository tables
  3. update() commits with the chunk transaction
  4. unique name prefixes context keys
  5. saveState=false → not restartable

basics

~20 s

ItemStream lets a reader/writer save and restore its position. open() initializes or restores state from the ExecutionContext, update() periodically writes the current position into it (at each chunk commit), and close() releases resources. On restart, open() reads the saved state so the step resumes where it stopped.

solid answer

~40 s

ItemStream is a callback interface (open, update, close) that gives stateful components a way to persist and recover their position via the step's ExecutionContext — a serializable key/value map Spring Batch stores in the JobRepository. open(ExecutionContext) is called once at step start: on a fresh run it initializes resources; on a restart it reads the previously saved keys and resumes at that point (e.g. line number, page, row count). update(ExecutionContext) is called at each chunk commit, writing the current position so a later failure has a checkpoint. close() releases resources (file handles, connections). Because update() runs inside the same commit as the writer, the saved position and the written data stay consistent. Most built-in readers implement ItemStream; a step wires them via CompositeItemStream. Setting saveState=false disables this, making the reader non-restartable.

code

java · 22 lines
java
public class ResumableReader extends AbstractItemCountingItemStreamItemReader<String> {
    private List<String> data;

    public ResumableReader() {
        setName("resumableReader"); // prefixes ExecutionContext keys
    }

    @Override protected void doOpen() {           // called from open(): acquire resource
        this.data = loadData();
    }

    @Override protected String doRead() {          // one item; base class tracks read count
        int idx = getCurrentItemCount() - 1;       // count already incremented
        return idx < data.size() ? data.get(idx) : null;
    }

    @Override protected void doClose() {           // release resource
        this.data = null;
    }
    // AbstractItemCountingItemStreamItemReader saves/restores currentItemCount
    // in the ExecutionContext automatically, so restart resumes at the right index.
}

go deeper

for a junior

Knows open/update/close exist and enable resume.

for a middle

Can explain ExecutionContext saves position for restart.

for a senior

Details commit-time update(), JobRepository persistence, saveState, naming, atomicity.

for a principal

Reasons about idempotency at chunk boundaries, custom ItemStream implementation, and restart correctness guarantees.

### The interface ```java public interface ItemStream { void open(ExecutionContext ctx) throws ItemStreamException; void update(ExecutionContext ctx) throws ItemStreamException; void close() throws ItemStreamException; } ``` `ItemStream` is how **stateful** step components (readers and writers) participate in **checkpointing and restart**. Many built-in readers implement `ItemStreamReader<T>` (which extends both `ItemReader<T>` and `ItemStream`). ### The ExecutionContext An **`ExecutionContext`** is a serializable **map of key/value pairs** attached to a step (and to the job) execution. Spring Batch persists it in the **`JobRepository`** (the batch metadata tables, e.g. `BATCH_STEP_EXECUTION_CONTEXT`). It is the durable 'memory' across a crash. There is one per step execution and one per job execution. ### Lifecycle within a step 1. **`open(ExecutionContext)`** — called **once** before reading begins. - **Fresh run:** the context is empty; the component initializes (opens the file, prepares the query) starting at the beginning. - **Restart:** the framework loads the *previous* execution's saved context, so `open()` finds keys like `"reader.read.count"` and **fast-forwards** to that position (e.g. skip N lines, jump to page P). 2. **`update(ExecutionContext)`** — called at **every chunk commit** (checkpoint). The component writes its **current** position into the context (`FlatFileItemReader` stores the current line count; paging readers store the page). This is the checkpoint that a later restart will read. 3. **`close()`** — called once when the step ends (success *or* failure) to release resources: close file handles, ResultSets, connections. ### Why the position stays consistent `update()` is invoked as part of the **chunk transaction commit**. So the persisted position and the actually-written output commit **together**. If the JVM dies mid-chunk, the transaction rolls back and the *last committed* position is the one saved — no items are silently lost or double-counted (assuming idempotent/transactional writing). ### Restart flow end-to-end 1. Job fails at chunk 50 of 100; the `ExecutionContext` in the JobRepository holds `read.count = 5000`. 2. Operator relaunches the **same job instance** (same identifying `JobParameters`). 3. Spring Batch finds the failed `StepExecution`, loads its saved `ExecutionContext`, and calls the reader's `open()` with it. 4. The reader reads `read.count`, skips forward to row 5000, and processing continues from there instead of the start. ### saveState and naming - Every stream needs a **unique `name`** (`setName(...)` / `.name(...)` in builders) — it prefixes the context keys so multiple streams don't collide. - **`saveState=true`** (default) enables persistence. Set **`saveState=false`** for stateless/throwaway readers or when you explicitly do *not* want restart to resume mid-stream (it will restart from scratch). - The step registers all streams via a `CompositeItemStream`, calling open/update/close on each in order. ### Gotchas - Reusing the same `name` across two streams → **key collisions**, corrupt restart. - Forgetting `saveState` implications: a non-restartable reader (e.g. reading from a queue) should set `saveState=false` to avoid a misleading resume. - **Restart requires the same JobInstance** — you must relaunch with the *same* identifying parameters; a new instance starts fresh. - Non-idempotent writers can double-process the boundary chunk unless the write is transactional/idempotent. - If you implement a **custom** reader that holds state, you must implement `ItemStream` yourself (or extend `AbstractItemCountingItemStreamItemReader`, which handles the read-count bookkeeping for you).

  • At what point during a step is update() called, and why does that timing matter?
    update() is called at each chunk commit, inside the same transaction that commits the chunk's writes. That timing makes the saved position and the written output atomic — a crash leaves the last committed position, so restart resumes cleanly without losing or duplicating committed items.
  • What is stored in the ExecutionContext for a FlatFileItemReader, and where does it persist?
    The current line/read count (keyed by the reader's name). It is serialized into the JobRepository's step execution context table (BATCH_STEP_EXECUTION_CONTEXT), so it survives a JVM crash and is available on restart.
  • What happens if you set saveState=false?
    The reader stops persisting its position to the ExecutionContext, so it is no longer restartable — a restart begins from the start of the stream. Use it for stateless sources or when mid-stream resume is undesirable.

saying these in an interview costs you the question

  • Thinking state is kept in memory rather than persisted to the JobRepository
  • Believing update() runs once at the end rather than per chunk commit
  • Assuming restart works with different (new) JobParameters / a new JobInstance
  • Forgetting streams need unique names, causing key collisions

context