Walk through what happens to JobInstance and JobExecution across a failure and a restart, including the metadata.
answer
- Fail -> ExecutionContext persisted, not lost
- Restart reuses instance, adds execution #2
- Restartable reader skips committed chunks
- 1 instance row, N execution rows in BATCH_ tables
- Different identity on restart = fresh instance, no resume
basics
~20 sThe first launch creates one JobInstance and a JobExecution that fails. On restart with the same identifying parameters, Spring Batch reuses the JobInstance, creates a second JobExecution, and resumes using the persisted ExecutionContext. Now the instance has two executions.
solid answer
~40 sLaunch one: JobLauncher asks JobRepository for the instance; none exists, so it creates a JobInstance (BATCH_JOB_INSTANCE) and a JobExecution (BATCH_JOB_EXECUTION). A step fails, the JobExecution ends FAILED, but its ExecutionContext (chunk position, custom state) is persisted. Restart: you relaunch with identical identifying parameters; the repository finds the existing, non-completed JobInstance and creates a second JobExecution linked to it. Restartable steps read the saved ExecutionContext and resume roughly where they stopped (e.g. skipping already-committed chunks), rather than starting from zero. When this attempt completes, the JobInstance is now COMPLETED and owns two JobExecutions. The metadata tables record the full history: one instance row, two execution rows, and StepExecution rows per attempt. This is why identity must stay stable across the failure/restart boundary.
code
java · 16 lines// A restartable chunk step: FlatFileItemReader saves its position in the
// step ExecutionContext after each commit, so a restart resumes mid-file.
Step loadStep = new StepBuilder("loadCustomers", jobRepository)
.<Customer, Customer>chunk(100, transactionManager)
.reader(flatFileReader) // implements ItemStream -> restartable
.writer(jdbcWriter)
.build();
Job job = new JobBuilder("customerLoad", jobRepository)
.start(loadStep)
.build();
// Launch #1 with runDate=2026-07-22 -> JobInstance + JobExecution #1, FAILS at line 5000.
// Relaunch with the SAME runDate=2026-07-22:
// -> same JobInstance, NEW JobExecution #2,
// -> reader reads saved context and resumes near line 5000, not line 1.go deeper
Only expected to know a restart adds an execution to the same instance.
Should mention the persisted ExecutionContext enabling resume.
Expected to trace the metadata tables and the restartable-component requirement precisely.
Should reason about restart guarantees, abandoned/start-limits, and designing readers/writers for exactly-once resume.
## The failure-and-restart lifecycle This is the scenario that makes the `JobInstance`/`JobExecution` split matter in practice. ### Step 0 — the model - **`JobInstance`** — logical run, identity = job name + identifying `JobParameters`. Row in **`BATCH_JOB_INSTANCE`**. - **`JobExecution`** — one attempt. Row in **`BATCH_JOB_EXECUTION`** (FK to instance). Owns `StepExecution`s (**`BATCH_STEP_EXECUTION`**) and two `ExecutionContext`s — job-level (**`BATCH_JOB_EXECUTION_CONTEXT`**) and step-level (**`BATCH_STEP_EXECUTION_CONTEXT`**). ### First attempt 1. `JobLauncher.run(job, params)` → `JobRepository` finds no matching `JobInstance` → creates one. 2. A new `JobExecution` is created (`BatchStatus.STARTING` → `STARTED`). 3. Steps run. In a chunk-oriented step, after each chunk commit the framework saves the reader/writer state into the step `ExecutionContext` (e.g. a `FlatFileItemReader` stores the current line count). 4. A step throws → its `StepExecution` and the `JobExecution` end **`FAILED`**. Crucially, the **persisted `ExecutionContext` survives** in the metadata tables. State now: **1 instance, 1 execution (FAILED)**. ### Restart 5. You relaunch with the **same identifying parameters**. The `JobRepository` finds the existing `JobInstance` (identity matches) and sees its last execution was not `COMPLETED`. 6. It creates a **new (second) `JobExecution`** under the same instance. 7. For **restartable** steps, Spring Batch loads the previously saved `ExecutionContext`. A restartable reader (like `FlatFileItemReader`) uses it to **skip already-processed items** and continue; already-committed chunks are not re-done. 8. Steps that already `COMPLETED` in the prior execution are, by default, **not re-run** (unless configured with `allowStartIfComplete(true)`). 9. This attempt succeeds → `JobExecution #2` and the `JobInstance` are now `COMPLETED`. State now: **1 instance, 2 executions** (FAILED, then COMPLETED). ### What the metadata looks like - `BATCH_JOB_INSTANCE`: **one** row (`JOB_INSTANCE_ID`, `JOB_NAME`, `JOB_KEY` hash of identifying params). - `BATCH_JOB_EXECUTION`: **two** rows, both pointing at that instance, with statuses `FAILED` and `COMPLETED`. - `BATCH_STEP_EXECUTION`: rows per step per execution. - Context tables hold the serialized `ExecutionContext` used to resume. ### Key rules and gotchas - **Identity must not change across the boundary.** If a restart uses different identifying parameters (e.g. a fresh timestamp), it creates a **new `JobInstance`** and starts from scratch — you lose resume-from-failure. - **Only restartable components resume.** Custom readers/writers must implement `ItemStream` and honor the `ExecutionContext` to truly resume; otherwise a restart re-processes from the top of the step. - **Completed instances can't restart.** Once the instance is `COMPLETED`, relaunching with the same identity throws `JobInstanceAlreadyCompleteException`. - **`ABANDONED`**: an execution can be marked abandoned so it is skipped when determining restartability. - **Restart limits**: `StepBuilder.startLimit(n)` caps how many times a step may be (re)started across executions. - **The `Step` scope of state**: `StepExecution` (and its counts — read/write/skip/commit) belong to each `JobExecution`, so per-attempt metrics are separate even though they roll up under one instance.
- During a restart, why does a plain in-memory reader re-process items from the start while a FlatFileItemReader does not?FlatFileItemReader implements ItemStream and writes its read position into the ExecutionContext on each commit, so on restart it restores that position. A component that ignores the ExecutionContext has no saved state to resume from, so it starts over.
- If the restart used a different identifying parameter than the failed run, what would happen to resume behavior?Identity would no longer match, so Spring Batch would create a brand-new JobInstance and JobExecution starting from scratch — the previous execution's ExecutionContext would not be consulted, and no resume would occur.
saying these in an interview costs you the question
- Saying restart starts a new JobInstance
- Claiming the ExecutionContext is discarded on failure
- Assuming every reader automatically resumes without implementing ItemStream
- Thinking StepExecution counts are shared across executions rather than per-attempt