Treating a Job as an ordered container of Steps, how do JobInstance/JobExecution/StepExecution identity and the JobRepository enable restart, and what are the design implications of decomposing a process into Steps?
answer
- JobInstance = name + identifying params
- JobExecution = one attempt; StepExecution = one step attempt
- JobRepository persists it -> restart possible
- Restart skips COMPLETED steps, restores ExecutionContext
- No job-wide tx; steps = restart/failure/commit seams
basics
~20 sA JobInstance is identified by job name + identifying parameters; each attempt is a JobExecution, each step run a StepExecution, all persisted in the JobRepository. On restart of the same instance, completed Steps are skipped so the job resumes mid-way — which is why splitting a process into Steps gives restart granularity.
solid answer
~50 sA Job being an ordered container of Steps only becomes powerful because of the runtime identity model in the JobRepository. A **JobInstance** = job name + *identifying* JobParameters; it's the logical 'this run of this input'. Each attempt at that instance is a **JobExecution**; each step attempt within it is a **StepExecution**, each with its own ExecutionContext. On failure the JobExecution and offending StepExecution are persisted FAILED. Restarting with the *same* identifying parameters resolves the same JobInstance; the SimpleJob's StepHandler skips Steps recorded COMPLETED and resumes at the failed one, restoring that step's ExecutionContext so a reader continues where it left off. Design implications: Step boundaries are your restart/commit granularity and failure-isolation seams; each step commits independently, so a mid-job crash leaves earlier steps' data committed. You choose step decomposition to bound rework, isolate failure domains, and expose operational checkpoints — while guarding against non-idempotent steps and duplicate identifying parameters (which would spawn a new instance or refuse to rerun).
code
java · 15 lines@Bean
public Job dailyImportJob(JobRepository jobRepository, Step extract, Step transform, Step load) {
return new JobBuilder("dailyImportJob", jobRepository)
// identifying param 'runDate' defines the JobInstance; a stable business key
// lets a failed run be RESTARTED (same instance), resuming at the failed step.
.incrementer(new RunIdIncrementer()) // adds non-identifying run.id for 'run again'
.start(extract)
.next(transform)
.next(load) // if 'load' fails, restart skips extract+transform (COMPLETED)
.build();
}
// Launch: identifying parameters resolve the JobInstance
// jobLauncher.run(dailyImportJob,
// new JobParametersBuilder().addString("runDate", "2026-07-22").toJobParameters());go deeper
Know a failed job can be restarted and completed steps are skipped.
Distinguish JobInstance/JobExecution/StepExecution and that the JobRepository persists them for restart.
Explain identifying parameters, ExecutionContext restore, allowStartIfComplete/startLimit, and per-step commit boundaries.
Reason about step decomposition as an architectural trade-off (restart granularity vs complexity), idempotency/side-effect design, parameter identity strategy, and repository durability requirements.
This is the 'why the container model matters' question. A `Job` is an ordered container of `Step`s, but the *value* of that structure comes from Spring Batch's **runtime metadata identity** and the **JobRepository** that persists it. **The identity trio:** - **`JobInstance`** — a *logical* run, identified by **job name + the identifying subset of `JobParameters`**. Parameters can be marked non-identifying (via the builder / `JobParameter(identifying=false)`); only identifying ones contribute to instance identity. Example: a `schedule.date=2026-07-22` parameter makes 'the July 22 import' one instance, distinct from July 23. - **`JobExecution`** — a single *attempt* to run a JobInstance. A failed-then-restarted instance has multiple JobExecutions but one JobInstance. Carries `BatchStatus`, `ExitStatus`, start/end times, and a job-level `ExecutionContext`. - **`StepExecution`** — a single attempt of one Step within a JobExecution; has its own `BatchStatus`, read/write/commit counts, and a **step-level `ExecutionContext`** where readers/writers persist their position. **The JobRepository** persists all three (in the `BATCH_JOB_INSTANCE`, `BATCH_JOB_EXECUTION`, `BATCH_STEP_EXECUTION`, and `*_EXECUTION_CONTEXT` tables, or an in-memory map for tests). This persistence is the linchpin of restart. **How restart works, mechanically:** 1. You relaunch with the **same identifying JobParameters**. The `JobLauncher`/`JobRepository` resolves the *existing* JobInstance rather than creating a new one. (Relaunching a **COMPLETED** instance with the same parameters throws `JobInstanceAlreadyCompleteException` — completed instances can't rerun; that's why you use a `JobParametersIncrementer` like `RunIdIncrementer` for 'run it again' semantics.) 2. A new `JobExecution` is created for the retry. 3. The `SimpleJob` iterates steps in order; its `StepHandler` checks the JobRepository: steps already COMPLETED in this JobInstance are **skipped** (unless `allowStartIfComplete`); it resumes at the first non-complete step. 4. For the resumed step, `AbstractStep` **restores the step's ExecutionContext**, so an `ItemStream` reader (e.g., a `FlatFileItemReader`) continues from the last committed position instead of re-reading from the top. **Design implications of decomposing into Steps** (the principal-level payoff): - **Restart/rework granularity.** Step boundaries are checkpoints. A long process split into extract → transform → load means a load failure doesn't force re-extraction. Fewer, bigger steps = coarser recovery; more, smaller steps = finer recovery but more metadata/complexity. - **Failure isolation.** Each step is an independent failure domain with its own status and exit code; you can attach different retry/skip policies and listeners per step. - **Transaction boundaries.** There is **no job-wide transaction**. Each step commits its own chunks. So a crash leaves prior steps' committed data in place — which is exactly what makes resuming meaningful, but means steps must be designed to tolerate 'previous steps already applied'. - **Idempotency & side effects.** Because a step can be re-run (restart, or `allowStartIfComplete`), non-idempotent steps risk duplicate side effects (double emails, double inserts). Principal-level design pushes side effects to idempotent operations or guards them with the ExecutionContext / natural keys. - **Operational visibility.** Per-step metadata (counts, durations, exit status) becomes your monitoring surface; decomposition improves observability and targeted alerting. **Gotchas:** 1. Adding an identifying timestamp parameter every run makes every launch a *new* JobInstance — restart of a failed run then can't find the old instance to resume. Use a stable identifying key (business date) plus a non-identifying run-id, or an incrementer, deliberately. 2. `allowStartIfComplete(true)` re-runs a completed step on every restart — great for idempotent validation/setup, harmful for data loads. 3. `startLimit` bounds restart attempts per step; a persistently failing step eventually throws `StartLimitExceededException`. 4. In-memory JobRepository (test/dev) loses metadata on JVM restart — restart-from-failure won't work across process restarts; production needs the persistent (JDBC) repository.
- You relaunch a COMPLETED JobInstance with identical identifying parameters. What happens and why?It throws JobInstanceAlreadyCompleteException — a completed instance cannot be re-run with the same identifying parameters. To run 'the same job again' you change an identifying parameter or use a JobParametersIncrementer (e.g. RunIdIncrementer) so a new JobInstance is created; restart-from-failure, by contrast, requires the SAME instance (i.e., not completed).
- How does the resumed step avoid re-reading data it already processed before the crash?AbstractStep restores that step's persisted ExecutionContext from the JobRepository; ItemStream readers/writers stored their position (e.g., line count / cursor) there, so they continue from the last committed point rather than the beginning.
- What's the trade-off in choosing many small steps versus a few large steps?Many small steps give finer restart granularity, failure isolation, and observability but add metadata volume and orchestration complexity; fewer large steps are simpler but force more rework on failure and coarser monitoring.
saying these in an interview costs you the question
- Believing the entire Job runs in one transaction that rolls back on failure
- Thinking every launch with a fresh timestamp param can still restart the previous failed run
- Assuming steps are automatically idempotent and safe to re-run
- Claiming an in-memory JobRepository supports restart across JVM restarts
- Confusing JobInstance (logical run) with JobExecution (single attempt)