skip to content

What is StepContribution and how do its counts get aggregated into the StepExecution?

level: seniorimportance: must knowfreq 40%

answer

  1. per-chunk staging buffer
  2. createStepContribution() -> increment -> apply()
  3. rollback => contribution discarded, totals clean
  4. apply() is synchronized (multi-threaded safety)
  5. commit/rollback counts live on StepExecution, not contribution

basics

~10 s

StepContribution is a per-chunk buffer that accumulates read/write/filter/skip counts while a chunk is processed. When the chunk commits, StepExecution.apply(contribution) merges those deltas into the step's running totals.

solid answer

~40 s

StepContribution is a lightweight staging object created from a StepExecution (stepExecution.createStepContribution()) for each chunk. As the chunk-oriented tasklet reads, processes and writes, it increments the contribution's counters (incrementReadCount, incrementWriteCount, incrementFilterCount, the skip counters, and an ExitStatus). It is deliberately kept separate from the StepExecution so that partial, uncommitted counts aren't reflected in the durable totals — if the chunk transaction rolls back, the contribution is discarded and nothing leaks into the StepExecution. On a successful commit, StepExecution.apply(StepContribution) is called: it adds the contribution's read/write/filter/skip deltas to the aggregate counts and folds in the ExitStatus. apply() is synchronized on the StepExecution, which is what makes count aggregation safe under multi-threaded steps where several chunks (each with its own contribution) finish concurrently. commitCount and rollbackCount are tracked directly on the StepExecution, not via the contribution.

code

java · 13 lines
java
// A custom Tasklet receives the StepContribution and feeds counts back
public class ArchiveTasklet implements Tasklet {
    @Override
    public RepeatStatus execute(StepContribution contribution,
                                ChunkContext chunkContext) {
        int moved = archiveOldRecords();
        contribution.incrementWriteCount(moved); // shows up in StepExecution.writeCount
        return RepeatStatus.FINISHED;
    }
}
// On commit the framework calls: stepExecution.apply(contribution)
//   -> synchronized merge of read/write/filter/skip deltas + ExitStatus
// commitCount / rollbackCount are incremented directly on the StepExecution.

go deeper

for a junior

Know StepContribution is a temporary buffer that later updates the StepExecution.

for a middle

Explain the create -> increment -> apply-on-commit lifecycle and that rollback discards it.

for a senior

Explain synchronized apply(), thread-per-chunk isolation, and that commit/rollback counts bypass the contribution.

for a principal

Reason about metadata consistency under multi-threaded/partitioned steps and why staging is required for transactional correctness of counters.

### The problem StepContribution solves A chunk-oriented step processes items in **chunks** wrapped in a transaction. While a chunk is in flight, we're counting items read/written/filtered — but the chunk might **roll back**. If we incremented the durable `StepExecution` counters directly and then rolled back, the counts would be wrong (and the StepExecution may already be persisted). We also may run chunks on **multiple threads**. Both problems are solved by a per-chunk staging buffer: `StepContribution`. ### Lifecycle 1. For each chunk, the framework calls `stepExecution.createStepContribution()` to get a **fresh** `StepContribution`. It carries a snapshot of the step's skip counts (for skip-limit checks) and starts its own read/write/filter deltas at 0. 2. During read/process/write, the tasklet increments the **contribution**, not the StepExecution: `incrementReadCount()`, `incrementWriteCount(int)`, `incrementFilterCount(int)`, `incrementReadSkipCount()`, `incrementWriteSkipCount()`, `incrementProcessSkipCount()`, and it can set an `ExitStatus`. 3. **On successful commit**, the framework calls `stepExecution.apply(contribution)`. That method is roughly: ```java public synchronized void apply(StepContribution c) { readSkipCount += c.getReadSkipCount(); writeSkipCount += c.getWriteSkipCount(); processSkipCount+= c.getProcessSkipCount(); filterCount += c.getFilterCount(); readCount += c.getReadCount(); writeCount += c.getWriteCount(); exitStatus = exitStatus.and(c.getExitStatus()); } ``` The running totals now include this chunk. 4. **On rollback**, the contribution is simply **not applied** (discarded), so the StepExecution's totals never saw the failed chunk's partial counts. `rollbackCount` is incremented **directly** on the StepExecution. ### commitCount / rollbackCount are special Note `apply()` does **not** touch commit/rollback counts. Those are maintained on the StepExecution itself (`incrementCommitCount()` after a commit; rollback count on rollback), because they describe transaction outcomes, not per-item deltas buffered in the contribution. ### Why apply() is synchronized In a **multi-threaded step** (`taskExecutor` on the step) many chunks run in parallel, each with its own StepContribution. They all merge into the **one shared** StepExecution. `apply()` being `synchronized` serializes those merges so the aggregate counts don't race. The contribution being thread-local per chunk means the hot per-item increments need no locking. ### Where you encounter it directly Inside a custom `Tasklet`, `execute(StepContribution contribution, ChunkContext chunkContext)` hands you the contribution — you call `contribution.incrementWriteCount(n)` so your tasklet's work shows up in the StepExecution counts. You rarely touch it in chunk steps (the framework drives it). ### Gotchas - Reading `stepExecution.getWriteCount()` **mid-chunk** won't include the current in-flight chunk — it's still in the contribution until commit. - Don't try to increment the StepExecution counters yourself; go through the contribution so rollback semantics and thread-safety hold. - A discarded (rolled-back) chunk's reads still 'happened' but aren't counted; after skip/retry re-processing, the surviving items are counted via a fresh contribution.

  • Why not just increment the StepExecution counters directly during the chunk?
    Because a chunk can roll back — directly-applied partial counts would then be wrong and possibly already persisted. The contribution stages the deltas and is only applied on commit, so rollbacks leave the totals clean. It also isolates per-thread counting in multi-threaded steps.
  • How does count aggregation stay correct when a step runs on multiple threads?
    Each chunk gets its own StepContribution (no shared per-item state), and StepExecution.apply(contribution) is synchronized, so the concurrent merges into the single shared StepExecution are serialized and don't lose updates.

saying these in an interview costs you the question

  • Thinking counts are written straight to StepExecution during the chunk
  • Believing apply() also bumps commitCount/rollbackCount
  • Not knowing apply() is synchronized / why that matters for multi-threaded steps

context