skip to content

Walk through the relationships between BATCH_JOB_INSTANCE, JOB_EXECUTION, STEP_EXECUTION, and EXECUTION_CONTEXT, and how they enable restart.

level: seniorimportance: should knowfreq 45%

answer

  1. Instance → Executions → StepExecutions (one-to-many each)
  2. Identity = name + JOB_KEY hash of identifying params
  3. Completed instance can't rerun → change a param
  4. Restart reuses instance, new execution, reloads ExecutionContext
  5. Completed steps skipped unless allowStartIfComplete

basics

~20 s

One JobInstance (name + identifying params) has many JobExecutions (one per attempt). Each JobExecution has many StepExecutions. ExecutionContext rows store saved state per job and per step. On restart, Spring reuses the instance and reloads the ExecutionContext to resume.

solid answer

~40 s

The tables form a hierarchy. `BATCH_JOB_INSTANCE` is the logical job, keyed by job name + `JOB_KEY` (a hash of identifying parameters). Each instance owns one or more `BATCH_JOB_EXECUTION` rows — one per run attempt, carrying status/exit-code/timestamps. Each job execution owns `BATCH_STEP_EXECUTION` rows, one per step, with commit/read/write/skip counts. `BATCH_JOB_EXECUTION_CONTEXT` and `BATCH_STEP_EXECUTION_CONTEXT` store the serialized `ExecutionContext` — arbitrary key/value state saved at chunk boundaries. On **restart**, Spring Batch finds the existing instance for the same parameters, refuses if it already completed, otherwise creates a *new* JobExecution but **copies the last failed execution's ExecutionContext** so readers/writers resume from their saved position (e.g. line count, cursor). Completed steps are skipped by default; the failed step re-runs from its last commit point.

code

java · 17 lines
java
// A restartable reader saves/loads position via ExecutionContext (ItemStream contract).
public class CountingReader implements ItemReader<String>, ItemStream {
    private int index = 0;
    private static final String KEY = "reader.index";

    @Override public void open(ExecutionContext ctx) {
        if (ctx.containsKey(KEY)) index = ctx.getInt(KEY); // resume on restart
    }
    @Override public void update(ExecutionContext ctx) {
        ctx.putInt(KEY, index); // persisted at each chunk commit -> STEP_EXECUTION_CONTEXT
    }
    @Override public String read() { return index < 100 ? "item-" + index++ : null; }
    @Override public void close() { }
}

// Force a new JobInstance each run by varying an identifying parameter:
// new JobParametersBuilder().addLong("run.id", System.nanoTime()).toJobParameters();

go deeper

for a junior

Know the one-to-many hierarchy and that ExecutionContext stores resume state.

for a middle

Explain instance identity via parameters and the completed-instance-can't-rerun rule.

for a senior

Detail the restart flow, ExecutionContext reload, step-skipping, and allowStartIfComplete/startLimit.

for a principal

Reason about restartability guarantees of stateful components and parameter-driven instance identity in scheduled/multi-node pipelines.

## The four-level hierarchy ``` BATCH_JOB_INSTANCE (logical job: name + JOB_KEY hash of identifying params) └── BATCH_JOB_EXECUTION (one per run ATTEMPT: status, start/end, exit code) ├── BATCH_STEP_EXECUTION (one per step: counts, status, exit) │ └── BATCH_STEP_EXECUTION_CONTEXT (serialized ExecutionContext for that step) └── BATCH_JOB_EXECUTION_CONTEXT (serialized ExecutionContext for the job) ``` Also: `BATCH_JOB_EXECUTION_PARAMS` holds the parameters of an execution, and sequences generate the IDs. ### JobInstance Identity = **job name + the identifying job parameters**. Spring hashes the identifying parameters into the `JOB_KEY` column. Two launches with the *same* name and identifying params refer to the **same** instance. Non-identifying parameters (marked `identifying=false`) do **not** change identity — handy for values like `run.date` you want to vary without creating a new instance, or the reverse. ### JobExecution A single **attempt** to run an instance. First run creates instance #1 and execution #1. If it fails and you relaunch with the same identifying params, you get the **same instance** but a **new execution** (#2). Status values: `STARTING`, `STARTED`, `COMPLETED`, `FAILED`, `STOPPED`, `ABANDONED`. ### StepExecution One row per step per job execution. Tracks `READ_COUNT`, `WRITE_COUNT`, `COMMIT_COUNT`, `ROLLBACK_COUNT`, `READ_SKIP_COUNT`, etc., plus status and exit status. This is what powers monitoring dashboards. ### ExecutionContext A persisted **key/value map** (`org.springframework.batch.item.ExecutionContext`) saved at **two scopes**: per job and per step. Stateful components — e.g. `FlatFileItemReader`, `JdbcCursorItemReader`, custom readers implementing `ItemStream` — write their progress here at chunk-commit boundaries (line number, last-read id). In Spring Batch 5 it is serialized as **JSON via Jackson** by default (older versions used Java serialization, a security/compat concern). ## How restart works 1. You relaunch a job with the **same identifying parameters**. 2. `JobRepository` finds the existing `JobInstance`. - If its last execution is `COMPLETED` → `JobInstanceAlreadyCompleteException` (a completed instance can't be rerun; change a parameter to make a new instance). - If currently running → `JobExecutionAlreadyRunningException`. - If `FAILED`/`STOPPED` → proceed. 3. A **new JobExecution** is created, but the **ExecutionContext from the last execution is loaded**, so stateful readers/writers resume from their saved position. 4. Steps already `COMPLETED` in the prior execution are **skipped** (default; unless the step is configured `allowStartIfComplete(true)`). The failed step restarts, resuming from its last committed chunk. ## Gotchas - To force a fresh run, change an identifying parameter (common pattern: a `run.id` incrementer) — otherwise a completed instance blocks re-execution. - `allowStartIfComplete(true)` and `startLimit(n)` alter default skip/limit behavior. - Restart correctness depends on readers/writers being **restartable** (`ItemStream` saving accurate state); a non-restartable custom reader will reprocess or skip data. - Because state is per-instance, sharing a metadata DB across environments/nodes must be coordinated (see isolation-on-create).

  • You relaunch a job with identical parameters after it completed successfully, and it throws JobInstanceAlreadyCompleteException. Why, and how do you run it again?
    A completed JobInstance cannot be re-executed. Change an identifying parameter — e.g. add an incrementing run.id (RunIdIncrementer) — so a new JobInstance is created.
  • On restart, why might a step still reprocess already-handled rows?
    Because its reader isn't truly restartable — it doesn't save/restore accurate position in the ExecutionContext via ItemStream — so it starts from the beginning of its (uncommitted) state instead of resuming.

saying these in an interview costs you the question

  • Saying a completed JobInstance can simply be rerun with the same parameters
  • Confusing JobInstance (logical) with JobExecution (attempt)
  • Thinking restart replays the whole job rather than resuming completed steps/last commit point
  • Believing ExecutionContext holds business data rather than reader/writer state

context