You inherit a nightly batch job that 'never resumes' after failures — it always reprocesses everything. Walk through the likely causes and how restart correctness depends on job/step design.
answer
- unique identifying param each run → new JobInstance → no resume
- reader needs ItemStream + saveState=true
- in-memory JobRepository loses state on crash
- multi-threaded step → use partitioning for restart
- check BATCH_JOB_INSTANCE / step ExecutionContext first
basics
~20 sMost often the job adds a unique identifying parameter (timestamp/UUID) each run, so every launch is a NEW JobInstance and there's nothing to resume. Other causes: readers with saveState=false or no ItemStream, an in-memory JobRepository losing state, or restartable=false.
solid answer
~40 sRestart depends on three things holding together. First, JobInstance identity: if the job appends a unique identifying parameter each run (a common 'always run' idiom), every launch is a fresh JobInstance and can never resume — fix by keeping identifying parameters stable and adding uniqueness only as non-identifying, or use a RunIdIncrementer deliberately. Second, per-step resume state: the reader must be stateful, implement ItemStream, and keep saveState=true; a custom reader without ItemStream, or saveState disabled for a multi-threaded step, silently reprocesses. Third, durable metadata: an in-memory/transient JobRepository loses all state on crash, so nothing survives to restart from. Also check the job's restartable flag and any allowStartIfComplete steps that re-run by design. Confirm the real symptom via the BATCH_JOB_INSTANCE / BATCH_STEP_EXECUTION_CONTEXT tables before changing code.
code
java · 12 lines// FIX for the most common cause: keep identity stable, uniqueness non-identifying.
JobParameters params = new JobParametersBuilder()
.addString("businessDate", "2026-07-22") // identifying -> stable across retries
.addLong("scheduledAt", epochMillis, false) // NON-identifying -> doesn't fork the instance
.toJobParameters();
// Retry after a FAILED run with the SAME businessDate => same JobInstance => resumes.
// Custom reader made restartable by implementing ItemStream via the base class:
public class KeyedReader extends AbstractItemCountingItemStreamItemReader<Rec> {
// update()/open() from the base persist/restore the read count in the
// Step ExecutionContext, so restart skips already-read items.
}go deeper
Recognize the symptom and that stable parameters matter.
Identify the unique-parameter trap and the reader/saveState requirement.
Systematically diagnose across identity, reader state, and repository durability; know partitioning for parallel restart.
Reason end-to-end about restart correctness: instance identity design, at-least-once idempotency, parallelism strategy, and when one-shot is the right policy.
## Framing 'Never resumes' is almost always a design/config problem, not a framework bug. Diagnose by inspecting the metadata tables and configuration in order. ## Cause 1 — new JobInstance every run (most common) A JobInstance = job name + **identifying** JobParameters. Teams frequently add a unique parameter each launch to bypass 'already complete' errors: ```java // Anti-pattern for restart: unique identifying param each run params.addLong("run.id", System.currentTimeMillis()); // identifying by default ``` Every launch is then a **distinct JobInstance** with no prior FAILED execution to resume — so it always starts from zero. Diagnosis: `BATCH_JOB_INSTANCE` shows a new row per launch. Fixes: - Keep identifying parameters **stable** across attempts (e.g. business date) so retries map to the same instance. - If you need uniqueness for scheduling but still want resume, make the extra parameter **non-identifying** (`addLong("ts", x, false)`), or drive reruns deliberately with a known incrementer only for genuinely new runs. Understand `RunIdIncrementer`/`JobParametersIncrementer` create *new* instances — good for 'next run', bad for 'resume this run'. ## Cause 2 — reader can't persist position Resume-within-a-step needs the reader to save a restart marker into the Step ExecutionContext via `ItemStream`: - **Custom reader not implementing ItemStream** → nothing is saved → restart re-reads from the top. Fix: implement `ItemStream` or extend `AbstractItemCountingItemStreamItemReader`. - **`saveState=false`** → framework won't persist the reader's position. Often set for multi-threaded steps; restore it or switch to partitioning. - **Multi-threaded step** with a single reader → a scalar line-count marker is ambiguous across threads, so restart is unreliable by design. Prefer **partitioning** (each partition = its own StepExecution/ExecutionContext) for scalable, restartable parallelism. ## Cause 3 — non-durable JobRepository Restart reads persisted metadata. A `ResourcelessJobRepository`/in-memory/`Map`-based repository (or a repository on an ephemeral DB) loses everything on a crash — after a hard failure there's no JobExecution to resume. Fix: use a real, durable database for the `BATCH_*` tables. ## Cause 4 — job/step flags - `restartable=false` (`preventRestart()`) → relaunch throws `JobRestartException`; the operator may then 'work around' it by changing parameters (→ Cause 1). Decide whether the job really should be one-shot. - Steps with `allowStartIfComplete(true)` re-run every restart by design — expected, not a bug, but explains 'it redid that step.' ## Cause 5 — idempotency masking / correctness Even when resume works, the failing chunk is re-read (at-least-once). If writes aren't idempotent you get duplicates that look like 'it reprocessed.' Design writers to be idempotent (upserts, dedupe keys) so re-run of the in-flight chunk is safe. ## Diagnostic checklist 1. `SELECT * FROM BATCH_JOB_INSTANCE` — one instance per business input, or a new one each launch? 2. Are JobParameters stable/identifying across attempts? 3. Reader: stateful + ItemStream + saveState=true? Multi-threaded? 4. JobRepository durable (real DB) and surviving crashes? 5. Job restartable flag; step allowStartIfComplete/startLimit settings. 6. Writes idempotent for the re-read chunk? ## When restart is the wrong tool For massively parallel or sharded loads, prefer **partitioning** (restartable per partition). For inherently non-idempotent one-shot operations, consider `preventRestart()` plus a fresh-instance rerun policy rather than pretending to resume.
- How would you make a scalable parallel load still restartable?Use partitioning: the master splits work into partitions, each running as its own StepExecution with its own ExecutionContext. On restart, completed partitions are skipped and failed ones resume independently — unlike a single multi-threaded reader whose positional marker is ambiguous across threads.
- The reader resumes correctly but you still see duplicate rows after a restart. Why?The chunk in flight at failure was rolled back and re-read on restart (at-least-once). If the writer isn't idempotent, those re-processed items are inserted twice. Fix with idempotent writes — upserts or a unique/dedupe key — not by changing restart config.
saying these in an interview costs you the question
- Blaming the framework instead of checking JobInstance identity / parameters
- Recommending a unique run parameter to 'fix' restart — that's what breaks it
- Assuming a multi-threaded step is safely restartable with a single positional reader
- Ignoring JobRepository durability when reasoning about crash recovery
- Expecting exactly-once writes for the in-flight chunk without designing idempotency