A cleanup Tasklet that returns CONTINUABLE to delete data in batches fails halfway. What determines whether the restarted job resumes correctly, and how do you design for it?
answer
- Each CONTINUABLE pass commits -> done work survives
- Restart re-runs execute() from the top, not mid-pass
- Idempotent/self-limiting query = easiest resume
- ExecutionContext marker committed with the tx
- COMPLETED steps need allowStartIfComplete to re-run
basics
~20 sBecause each CONTINUABLE pass commits its own transaction, completed batches stay deleted. On restart the whole step re-runs execute() from the start, so correctness depends on the work being idempotent or on progress saved in the ExecutionContext.
solid answer
~50 sWith a CONTINUABLE Tasklet, every pass runs and commits in its own transaction, so batches completed before the crash remain committed. On restart, Spring Batch re-executes the same step and calls execute() from the beginning — it does not resume mid-execute. So correct resumption depends on one of two things: either the operation is naturally idempotent/self-limiting (e.g. DELETE ... WHERE created < cutoff LIMIT n — already-deleted rows can't match again, so re-running simply continues), or you persist a progress marker (offset, last-id) in the StepExecution's ExecutionContext, which the JobRepository saves on each commit; the restarted step reads it back and continues. Pitfalls: non-idempotent work double-applies; a job marked COMPLETED won't restart at all; and if a completed step must re-run you need allowStartIfComplete. Design the delete to be replayable rather than relying on saved cursors when possible.
code
java · 17 lines@Bean
@StepScope
Tasklet purgeTasklet(JdbcTemplate jdbc) {
return (contribution, chunkContext) -> {
// Self-limiting: already-deleted rows no longer match, so restart resumes naturally
int n = jdbc.update(
"DELETE FROM audit_log WHERE created < now() - interval '90 days' LIMIT 50000");
contribution.incrementWriteCount(n);
// Optional progress marker, committed atomically with the delete in this tx
ExecutionContext ec = chunkContext.getStepContext()
.getStepExecution().getExecutionContext();
ec.putLong("purged", ec.getLong("purged", 0) + n);
return RepeatStatus.continueIf(n > 0);
};
}go deeper
Likely only knows committed work survives a crash.
Should know restart re-runs the step and that ExecutionContext can hold progress.
Should design idempotent/self-limiting deletes and persist markers in ExecutionContext.
Should reason about transactional atomicity of the marker, offset drift, side-effect idempotency, restartable state/job-parameter identity, and prefer replayable operations over saved cursors.
## The scenario A Tasklet deletes rows in batches, returning `RepeatStatus.CONTINUABLE` each pass until nothing remains. Halfway through, the job crashes. What guarantees a clean restart? ## What Spring Batch actually does on restart 1. **Each CONTINUABLE pass is its own transaction.** The framework commits after each `execute()` that returns CONTINUABLE. So all batches deleted before the crash are **already committed** — they won't come back. 2. **Restart re-runs the step from the top.** Spring Batch does not snapshot the middle of an `execute()` call. When you relaunch the same job (same `JobInstance` / identifying job parameters), a failed step is re-executed: `execute()` is called again from the beginning. There is no "resume at pass N" magic — resumption is *your* responsibility via idempotency or saved state. 3. **JobRepository holds the truth.** The `JobRepository` persists `StepExecution` status and its `ExecutionContext` at each commit. A step that ended `FAILED` is eligible for re-execution; a step/job that ended `COMPLETED` is **not** restartable unless you set `allowStartIfComplete(true)`. ## Two ways to be restart-correct ### 1. Idempotent / self-limiting work (preferred) Design the operation so replaying it is harmless: ```java int n = jdbc.update("DELETE FROM audit WHERE created < :cutoff AND id IN (SELECT id ... LIMIT 50000)"); return RepeatStatus.continueIf(n > 0); ``` Already-deleted rows no longer satisfy the predicate, so a restart naturally continues from where the data now stands — no saved cursor needed. This is the most robust approach because it doesn't depend on any persisted position being consistent with the actual data. ### 2. Persisted progress in ExecutionContext When the work isn't self-limiting (e.g. paging by key over stable data), store the marker: ```java ExecutionContext ec = chunkContext.getStepContext() .getStepExecution().getExecutionContext(); long lastId = ec.getLong("lastId", 0L); long newLast = processBatchAfter(lastId); ec.putLong("lastId", newLast); return RepeatStatus.continueIf(newLast > lastId); ``` The `JobRepository` saves the `ExecutionContext` on commit, so a restarted step reads back `lastId` and resumes. **Caveat**: the saved marker must remain valid against the data — a numeric offset over a mutating table can skip or double-process rows; prefer a stable key (last processed id) over a positional offset. ## Gotchas at principal level - **Double-application on non-idempotent steps** (incrementing counters, sending emails, appending files) — CONTINUABLE + restart can repeat side effects for the pass that was in flight at crash time. Make side effects idempotent or transactional with the progress marker in the *same* transaction. - **COMPLETED steps don't restart.** If the step finished but a later step failed, re-running won't re-invoke the completed cleanup unless `allowStartIfComplete(true)` — usually you *don't* want cleanup to re-run, so leave it. - **Transaction scope of the marker.** Because the ExecutionContext is committed with the step's transaction, the progress marker and the deleted rows commit atomically — that's what makes marker-based resume safe. If you did the delete outside the managed transaction, the marker and data could diverge. - **Job parameters identity.** Restart must use the same identifying job parameters, or Spring Batch creates a *new* JobInstance and starts fresh rather than resuming. - **Idempotency beats bookkeeping.** Whenever feasible, make the operation replayable (approach 1) so restart correctness doesn't hinge on a saved cursor staying consistent. ## Bottom line Correct resumption is determined by (a) committed passes surviving, (b) the step being in a restartable (FAILED) state with matching job parameters, and (c) your logic being idempotent or reading persisted ExecutionContext state. Design the delete to be self-limiting first; fall back to ExecutionContext markers committed in the same transaction when it can't be.
- Why is a self-limiting DELETE safer than paging by numeric offset for restart?A self-limiting DELETE re-evaluates the current data each pass, so replaying it can't skip or duplicate rows. A numeric offset over a mutating table can drift as rows are deleted, causing skips or reprocessing; a stable last-id marker is the middle ground.
- Where is the ExecutionContext stored and when is it persisted?In the batch metadata tables via the JobRepository, persisted at each step-transaction commit — so it commits atomically with the batch's data changes, keeping the marker consistent with what was actually done.
- If the cleanup step finished COMPLETED but a later step failed, will restart re-run the cleanup?No — a COMPLETED step is skipped on restart unless it was configured with allowStartIfComplete(true). Usually you leave cleanup non-repeating so it doesn't run twice.
saying these in an interview costs you the question
- Thinking Spring Batch resumes in the middle of a single execute() call
- Assuming CONTINUABLE work is automatically restart-safe with no design effort
- Relying on a numeric offset over mutating data without considering drift
- Not realizing COMPLETED steps won't re-run without allowStartIfComplete
- Doing the delete outside the managed transaction so the marker and data can diverge