How does the ExecutionContext enable a reader to resume mid-step after a failure? Walk through the ItemStream mechanics.
answer
- ItemStream: open/update/close
- update() per chunk writes read.count
- open() on restart fast-forwards (jumpToItem)
- AbstractItemCountingItemStreamItemReader + setName namespacing
- saveState=false disables resume
basics
~20 sStateful readers implement ItemStream. On open() they read a saved position from the ExecutionContext; on update() (called per chunk) they write the current read count. After a crash, restart calls open() with the saved context so the reader skips already-read items.
solid answer
~40 sRestartable readers implement the ItemStream interface: open(ExecutionContext), update(ExecutionContext), and close(). During a normal run, after each chunk Spring calls update(), where the reader stores its progress — for count-based readers that's the number of items already read, saved under a namespaced key. This context is committed with the chunk. If the step fails and you restart the same JobInstance, the framework reloads the step's persisted ExecutionContext and calls open() with it; the reader (e.g. FlatFileItemReader via AbstractItemCountingItemStreamItemReader) reads that saved count and fast-forwards past already-processed items so it resumes exactly where it stopped. Keys are namespaced by the reader's name (setName) to avoid collisions when multiple streams live in one step. You can disable this with saveState=false if you don't want restart tracking.
code
java · 25 linespublic class OffsetReader extends AbstractItemStreamItemReader<Record> {
private static final String KEY = "offset";
private long offset = 0;
private final DataSource source;
public OffsetReader(DataSource source) { this.source = source; setName("offsetReader"); }
@Override
public void open(ExecutionContext ctx) {
// resume point if restarting; getExecutionContextKey prefixes with name
this.offset = ctx.getLong(getExecutionContextKey(KEY), 0L);
}
@Override
public Record read() {
Record r = source.fetch(offset);
if (r != null) offset++;
return r; // null signals end of input
}
@Override
public void update(ExecutionContext ctx) { // called each chunk commit
ctx.putLong(getExecutionContextKey(KEY), offset);
}
}go deeper
Likely unaware of ItemStream; may just say 'it resumes'.
Knows open/update but may miss the transactional/per-chunk timing.
Should articulate the full open/update/close flow and namespacing.
Should surface ordering/determinism hazards for cursor/paging readers and custom-reader design.
## The ItemStream contract Restartability hangs on the `org.springframework.batch.item.ItemStream` interface, with three callbacks: - `open(ExecutionContext ctx)` — called once when the step starts (or restarts). The stream **reads** any previously saved state from `ctx` and positions itself. - `update(ExecutionContext ctx)` — called **once per chunk**. The stream **writes** its current progress into `ctx`. - `close()` — releases resources. Chunk-oriented steps register their reader/writer/processor streams; if a component implements `ItemStream`, the step drives these callbacks automatically. (`ItemStreamReader`/`ItemStreamWriter` are the reader/writer flavors.) ## Count-based readers Most built-in readers extend `AbstractItemCountingItemStreamItemReader`. It tracks a **current item count** and, in `update()`, does roughly `ctx.putInt(getExecutionContextKey("read.count"), currentItemCount)`. `getExecutionContextKey(...)` prefixes the key with the reader's **name** (set via `setName(...)`, backed by `ExecutionContextUserSupport`) so two readers in the same step don't clobber each other. On `open()`, if `isSaveState()` is true and the context contains the saved count, the reader sets its position and, for something like `FlatFileItemReader`, **reads and discards** that many lines to fast-forward to the resume point (`jumpToItem`). ## End-to-end restart flow 1. Run processes chunks; each commit persists `read.count = N` in `BATCH_STEP_EXECUTION_CONTEXT`, transactionally with the written data. 2. JVM crashes after chunk N. 3. Operator relaunches the **same JobInstance** (identical identifying parameters). 4. Spring finds the failed StepExecution, loads its saved context, creates a new StepExecution seeded with it, calls `open(ctx)`. 5. Reader reads `read.count = N`, skips the first N items, resumes at item N+1 — no gap, no duplicate (given a transactional writer). ## Key details & gotchas - **`saveState`**: set `saveState=false` to turn off context tracking (e.g. for a non-restartable reader or to reduce overhead) — but then it won't resume. - **Namespacing**: always give each stateful reader a unique `name`; duplicate names cause key collisions and corrupt resume state. - **Non-transactional resources**: fast-forwarding assumes deterministic ordering. A `JdbcCursorItemReader` restarts by re-running the query and skipping rows — if the underlying data or ordering changed between runs, resume can be wrong. `JdbcPagingItemReader` re-issues page queries; ordering must be stable and by a unique sort key. - **Custom readers**: implement `ItemStream` yourself and store just enough (an offset, a last-id) — keep it small and serializable. - **Processors/writers** can also be streams if they hold resumable state, but readers are the usual case.
- Why do reader keys get prefixed with the reader's name?So multiple stateful streams in the same step don't collide on the same key in the shared step ExecutionContext. setName / ExecutionContextUserSupport builds the namespaced key.
- What's a risk when a JdbcCursorItemReader resumes after a restart?It re-runs the query and skips already-read rows by count. If the row set or ordering changed between runs, it can skip or duplicate rows — resume assumes a stable, deterministic ordering.
saying these in an interview costs you the question
- Claiming the reader is stateless and Spring 'just knows' where to resume
- Forgetting update() runs per chunk and open() reloads on restart
- Not namespacing keys, causing collisions between readers