skip to content

A cleanup Tasklet that returns CONTINUABLE to delete data in batches fails halfway. What determines whether the restarted job resumes correctly, and how do you design for it?

level: principalimportance: should knowfreq 20%

answer

  1. Each CONTINUABLE pass commits -> done work survives
  2. Restart re-runs execute() from the top, not mid-pass
  3. Idempotent/self-limiting query = easiest resume
  4. ExecutionContext marker committed with the tx
  5. COMPLETED steps need allowStartIfComplete to re-run

basics

~20 s

Because each CONTINUABLE pass commits its own transaction, completed batches stay deleted. On restart the whole step re-runs execute() from the start, so correctness depends on the work being idempotent or on progress saved in the ExecutionContext.

solid answer

~50 s

With a CONTINUABLE Tasklet, every pass runs and commits in its own transaction, so batches completed before the crash remain committed. On restart, Spring Batch re-executes the same step and calls execute() from the beginning — it does not resume mid-execute. So correct resumption depends on one of two things: either the operation is naturally idempotent/self-limiting (e.g. DELETE ... WHERE created < cutoff LIMIT n — already-deleted rows can't match again, so re-running simply continues), or you persist a progress marker (offset, last-id) in the StepExecution's ExecutionContext, which the JobRepository saves on each commit; the restarted step reads it back and continues. Pitfalls: non-idempotent work double-applies; a job marked COMPLETED won't restart at all; and if a completed step must re-run you need allowStartIfComplete. Design the delete to be replayable rather than relying on saved cursors when possible.

code

java · 17 lines
java
@Bean
@StepScope
Tasklet purgeTasklet(JdbcTemplate jdbc) {
    return (contribution, chunkContext) -> {
        // Self-limiting: already-deleted rows no longer match, so restart resumes naturally
        int n = jdbc.update(
            "DELETE FROM audit_log WHERE created < now() - interval '90 days' LIMIT 50000");
        contribution.incrementWriteCount(n);

        // Optional progress marker, committed atomically with the delete in this tx
        ExecutionContext ec = chunkContext.getStepContext()
            .getStepExecution().getExecutionContext();
        ec.putLong("purged", ec.getLong("purged", 0) + n);

        return RepeatStatus.continueIf(n > 0);
    };
}

go deeper

for a junior

Likely only knows committed work survives a crash.

for a middle

Should know restart re-runs the step and that ExecutionContext can hold progress.

for a senior

Should design idempotent/self-limiting deletes and persist markers in ExecutionContext.

for a principal

Should reason about transactional atomicity of the marker, offset drift, side-effect idempotency, restartable state/job-parameter identity, and prefer replayable operations over saved cursors.

## The scenario A Tasklet deletes rows in batches, returning `RepeatStatus.CONTINUABLE` each pass until nothing remains. Halfway through, the job crashes. What guarantees a clean restart? ## What Spring Batch actually does on restart 1. **Each CONTINUABLE pass is its own transaction.** The framework commits after each `execute()` that returns CONTINUABLE. So all batches deleted before the crash are **already committed** — they won't come back. 2. **Restart re-runs the step from the top.** Spring Batch does not snapshot the middle of an `execute()` call. When you relaunch the same job (same `JobInstance` / identifying job parameters), a failed step is re-executed: `execute()` is called again from the beginning. There is no "resume at pass N" magic — resumption is *your* responsibility via idempotency or saved state. 3. **JobRepository holds the truth.** The `JobRepository` persists `StepExecution` status and its `ExecutionContext` at each commit. A step that ended `FAILED` is eligible for re-execution; a step/job that ended `COMPLETED` is **not** restartable unless you set `allowStartIfComplete(true)`. ## Two ways to be restart-correct ### 1. Idempotent / self-limiting work (preferred) Design the operation so replaying it is harmless: ```java int n = jdbc.update("DELETE FROM audit WHERE created < :cutoff AND id IN (SELECT id ... LIMIT 50000)"); return RepeatStatus.continueIf(n > 0); ``` Already-deleted rows no longer satisfy the predicate, so a restart naturally continues from where the data now stands — no saved cursor needed. This is the most robust approach because it doesn't depend on any persisted position being consistent with the actual data. ### 2. Persisted progress in ExecutionContext When the work isn't self-limiting (e.g. paging by key over stable data), store the marker: ```java ExecutionContext ec = chunkContext.getStepContext() .getStepExecution().getExecutionContext(); long lastId = ec.getLong("lastId", 0L); long newLast = processBatchAfter(lastId); ec.putLong("lastId", newLast); return RepeatStatus.continueIf(newLast > lastId); ``` The `JobRepository` saves the `ExecutionContext` on commit, so a restarted step reads back `lastId` and resumes. **Caveat**: the saved marker must remain valid against the data — a numeric offset over a mutating table can skip or double-process rows; prefer a stable key (last processed id) over a positional offset. ## Gotchas at principal level - **Double-application on non-idempotent steps** (incrementing counters, sending emails, appending files) — CONTINUABLE + restart can repeat side effects for the pass that was in flight at crash time. Make side effects idempotent or transactional with the progress marker in the *same* transaction. - **COMPLETED steps don't restart.** If the step finished but a later step failed, re-running won't re-invoke the completed cleanup unless `allowStartIfComplete(true)` — usually you *don't* want cleanup to re-run, so leave it. - **Transaction scope of the marker.** Because the ExecutionContext is committed with the step's transaction, the progress marker and the deleted rows commit atomically — that's what makes marker-based resume safe. If you did the delete outside the managed transaction, the marker and data could diverge. - **Job parameters identity.** Restart must use the same identifying job parameters, or Spring Batch creates a *new* JobInstance and starts fresh rather than resuming. - **Idempotency beats bookkeeping.** Whenever feasible, make the operation replayable (approach 1) so restart correctness doesn't hinge on a saved cursor staying consistent. ## Bottom line Correct resumption is determined by (a) committed passes surviving, (b) the step being in a restartable (FAILED) state with matching job parameters, and (c) your logic being idempotent or reading persisted ExecutionContext state. Design the delete to be self-limiting first; fall back to ExecutionContext markers committed in the same transaction when it can't be.

  • Why is a self-limiting DELETE safer than paging by numeric offset for restart?
    A self-limiting DELETE re-evaluates the current data each pass, so replaying it can't skip or duplicate rows. A numeric offset over a mutating table can drift as rows are deleted, causing skips or reprocessing; a stable last-id marker is the middle ground.
  • Where is the ExecutionContext stored and when is it persisted?
    In the batch metadata tables via the JobRepository, persisted at each step-transaction commit — so it commits atomically with the batch's data changes, keeping the marker consistent with what was actually done.
  • If the cleanup step finished COMPLETED but a later step failed, will restart re-run the cleanup?
    No — a COMPLETED step is skipped on restart unless it was configured with allowStartIfComplete(true). Usually you leave cleanup non-repeating so it doesn't run twice.

saying these in an interview costs you the question

  • Thinking Spring Batch resumes in the middle of a single execute() call
  • Assuming CONTINUABLE work is automatically restart-safe with no design effort
  • Relying on a numeric offset over mutating data without considering drift
  • Not realizing COMPLETED steps won't re-run without allowStartIfComplete
  • Doing the delete outside the managed transaction so the marker and data can diverge

context