skip to content

What are the ChunkListener callbacks, and how do they relate to the chunk's transaction boundaries?

level: seniorimportance: should knowfreq 40%

answer

  1. beforeChunk / afterChunk / afterChunkError
  2. chunk = transaction unit
  3. beforeChunk inside tx, afterChunk after commit
  4. afterChunkError after rollback
  5. may repeat across retries; ChunkContext -> StepExecution

basics

~10 s

ChunkListener has beforeChunk, afterChunk, and afterChunkError. beforeChunk runs at the start of the chunk inside its transaction; afterChunk runs after that transaction commits successfully; afterChunkError runs if the chunk failed and rolled back.

solid answer

~40 s

A chunk is the unit of transaction in a chunk-oriented step. ChunkListener brackets that unit: beforeChunk(ChunkContext) runs after the transaction is started but before reading/processing/writing; afterChunk(ChunkContext) runs after the chunk's transaction has committed; afterChunkError(ChunkContext) runs when the chunk fails, after the rollback. The transactional placement matters: work in beforeChunk participates in the chunk transaction and is rolled back with it, whereas afterChunk runs post-commit, so anything you do there is outside that transaction. Because afterChunk only runs on success and afterChunkError only on failure, they're mutually exclusive per chunk attempt. ChunkContext gives access to the StepContext (hence StepExecution and ExecutionContext) and a rollback flag. Annotation forms: @BeforeChunk, @AfterChunk, @AfterChunkError. This is the right scope for chunk-level counters, flushing side buffers post-commit, or logging failed-chunk diagnostics.

code

java · 18 lines
java
public class ChunkMetricsListener implements ChunkListener {
    @Override
    public void beforeChunk(ChunkContext ctx) {
        // inside the chunk transaction — rolled back if the chunk fails
    }
    @Override
    public void afterChunk(ChunkContext ctx) {
        // AFTER commit — data is durable; do post-commit side effects here
        StepExecution se = ctx.getStepContext().getStepExecution();
        progress.set(se.getWriteCount());
    }
    @Override
    public void afterChunkError(ChunkContext ctx) {
        // AFTER rollback — log the failed chunk
        log.warn("Chunk failed at write count {}",
                 ctx.getStepContext().getStepExecution().getWriteCount());
    }
}

go deeper

for a junior

Name the three callbacks and that a chunk is a transaction.

for a middle

Know afterChunk = success/after commit and afterChunkError = failure/after rollback.

for a senior

Explain that beforeChunk is transactional but afterChunk is post-commit, and the consequences for side effects.

for a principal

Reason about at-least-once semantics under retry, idempotent post-commit side effects, and why throwing from afterChunk is hazardous.

## Chunk = the transaction boundary In a chunk-oriented step the framework: starts a transaction, reads N items (the chunk size), processes them, writes them, and commits. If anything throws, the transaction **rolls back**. `ChunkListener` wraps this transactional unit. ## The three callbacks - `void beforeChunk(ChunkContext context)` — called **after the transaction has begun** but before the read/process/write cycle for that chunk. Work here runs *inside* the chunk transaction. - `void afterChunk(ChunkContext context)` — called **after the chunk transaction commits**. It runs only on success; work here is *outside* the just-committed transaction. - `void afterChunkError(ChunkContext context)` — called when the chunk **fails**, after the rollback. Only on failure. `afterChunk` and `afterChunkError` are mutually exclusive for a given chunk attempt: success -> afterChunk, failure -> afterChunkError. ## Why the transaction placement matters This is the crux and a favorite interview probe: - **beforeChunk is transactional.** If you write to the database in beforeChunk and the chunk later fails, that write is rolled back along with the chunk. Good for setup that should be atomic with the chunk. - **afterChunk is post-commit.** The chunk's data is already durable. Anything you do here is in a *new* transaction context (or none). Don't assume you can "undo" it by failing the chunk — the chunk already committed. Also, an exception thrown from afterChunk is problematic because the commit already happened. - **afterChunkError is post-rollback.** Use it for diagnostics/logging of the failed chunk. Note you generally can't rely on the DB state being anything, since it rolled back. ## ChunkContext `ChunkContext` provides: - `getStepContext()` -> `StepContext` -> `StepExecution`, `JobExecution`, `JobParameters`, and both step and job `ExecutionContext`s. - `isComplete()` / attribute bag and a rollback indicator used internally. ## Retry interaction In a fault-tolerant step, a failed chunk may be **retried** (retry is a sibling leaf, but the listener consequence is worth knowing): the chunk is re-attempted, so `beforeChunk`/`afterChunkError` can fire multiple times for the same logical chunk across attempts. Don't treat these callbacks as exactly-once. ## Annotation forms `@BeforeChunk`, `@AfterChunk`, `@AfterChunkError` — on any bean, registered via `StepBuilder.listener(...)`. ## Registration ```java new StepBuilder("s", jobRepository) .<In, Out>chunk(100, transactionManager) .reader(r).processor(p).writer(w) .listener(chunkListener) .build(); ``` ## When to use - Maintain chunk-level metrics/progress. - Post-commit side effects that must only happen once the data is durable (e.g., enqueue a message) — in afterChunk. - Setup that must be atomic with the chunk — in beforeChunk. - Failed-chunk logging / alerting — in afterChunkError. ## Gotchas - Throwing from afterChunk is dangerous: the chunk is already committed, so the exception can't roll it back and may fail the step in a confusing state. - afterChunk is NOT a rollback-safe place to write data you might want to undo. - With single-item chunks (chunk size 1) these fire per item, which can be expensive. - Don't confuse afterChunkError with SkipListener/RetryListener — afterChunkError fires on chunk failure regardless of whether items are ultimately skipped.

  • If you insert an audit row in afterChunk and the next chunk fails, is the audit row rolled back?
    No. afterChunk runs after the current chunk committed, in a separate transaction context, and the next chunk's failure is a different transaction. The audit row from afterChunk is not part of either and won't be undone by the later failure.
  • Can beforeChunk and afterChunkError both fire for the same chunk?
    Yes. beforeChunk runs at the start of the attempt; if that attempt then fails, afterChunkError runs. afterChunk (success) is the one that's mutually exclusive with afterChunkError.

saying these in an interview costs you the question

  • Saying afterChunk runs inside the chunk transaction
  • Thinking work in afterChunk can be rolled back by failing the chunk
  • Assuming each callback fires exactly once even under retry
  • Confusing afterChunkError with the skip/retry listeners

context