Why did Spring Batch split the run concept into JobInstance and JobExecution instead of a single object, and what does that buy an operator?
answer
- Logical once-only vs physical retry — two needs
- Instance = idempotency unit; Execution = retry/observability unit
- Identity as persisted key -> AlreadyComplete guard
- Identifying-param choice = idempotency boundary
- Metadata store = source of truth, must be durable
basics
~10 sSeparating 'the logical run' (JobInstance) from 'each attempt' (JobExecution) lets Spring Batch guarantee a run happens once while still allowing many retry attempts. It enables restartability, idempotency, per-attempt metrics, and a clean audit history.
solid answer
~40 sThe split encodes a fundamental batch requirement: a logical unit of work (say end-of-day for a date) must run to success exactly once, yet may need several physical attempts. Modeling identity separately (JobInstance = job name + identifying parameters) gives the framework a stable key to (a) refuse duplicate completed runs via JobInstanceAlreadyCompleteException, guaranteeing idempotency; (b) locate a failed instance and restart it, resuming from the persisted ExecutionContext; and (c) attach each attempt's runtime state, status, timings and StepExecution metrics to its own JobExecution while still grouping them under one instance for audit. Operationally you get: safe scheduler reruns, resumable long jobs, per-attempt observability, and a queryable history in the metadata tables. A single combined object couldn't express 'same logical run, different attempt', which is exactly what restart and idempotency depend on.
go deeper
Not expected at this depth; would just restate the basic distinction.
Can list benefits (restart, audit) but usually not the idempotency-boundary reasoning.
Should connect the split to restartability and per-attempt metrics.
Expected to reason about parameter/identity design as an idempotency contract, metadata durability, and the counterfactual of a single object.
## Why two objects instead of one The `JobInstance`/`JobExecution` separation is not incidental — it is the data model that makes Spring Batch's core guarantees expressible. ### The requirement being modeled Batch workloads have a tension: - A **logical unit of work** must be done **once and only once** — 'settle trades for 2026-07-22' should not be applied twice. - But the **physical execution** may **fail and be retried** many times (transient DB outages, node crashes, deploys mid-run). A single object can represent *either* 'the logical run' *or* 'one attempt', but not both cleanly. Spring Batch splits them: - **`JobInstance`** = the logical run, keyed by **job name + identifying `JobParameters`**. It is the unit of *idempotency*. - **`JobExecution`** = one physical attempt, holding status, timing, `ExecutionContext`, and `StepExecution`s. It is the unit of *retry and observability*. ### What the split buys you **1. Idempotency guarantee.** Because identity is a first-class, persisted key, the `JobRepository` can enforce that a `COMPLETED` `JobInstance` is never re-run with the same identity — it throws **`JobInstanceAlreadyCompleteException`**. A scheduler that fires twice, or a manual rerun after a network blip, cannot double-apply the work. This is far stronger than best-effort application-level dedup. **2. Restartability.** The stable instance identity lets the framework find a prior failed attempt and its persisted `ExecutionContext`, then create a fresh `JobExecution` that resumes. Without a separate 'attempt' object, you could not represent 'attempt #2 of the same logical run'. **3. Per-attempt observability, instance-level rollup.** Each `JobExecution` carries its own `BatchStatus`, `ExitStatus`, start/end times, and `StepExecution` read/write/skip/commit counts. You can see *why attempt #1 failed* separately from *how attempt #2 succeeded*, yet both roll up under one instance for 'was this run done?'. **4. Clean audit history.** The metadata schema (`BATCH_JOB_INSTANCE`, `BATCH_JOB_EXECUTION`, `BATCH_STEP_EXECUTION`, plus context tables) is queryable: operators can answer 'how many times did the 07-22 run attempt before succeeding?' with a join, not log spelunking. ### Design and operational consequences - **Parameter design is an architectural decision.** What you mark as identifying defines the idempotency boundary. Put the business date in identity → one run per date (restartable). Add a timestamp to identity → every launch is unique (never restartable). Principals choose this deliberately, often via a `JobParametersIncrementer` (`RunIdIncrementer`) for jobs that *should* be new each time. - **Restart contracts propagate to components.** Readers/writers must implement `ItemStream` and honor the `ExecutionContext` to make the resume guarantee real; the domain model enables restart but the components must cooperate. - **Metadata store is shared infrastructure.** The `JobRepository` tables are the source of truth for idempotency and restart; they need a real, durable, transactional datastore in production (not an in-memory map), and their transactional boundaries matter for exactly-once semantics. - **`ABANDONED` / start limits.** Operators can abandon a stuck execution so it is skipped in restart evaluation, and cap re-starts via `startLimit` — controls that only make sense because attempts are modeled separately from the instance. ### The counterfactual If Spring Batch had a single 'JobRun' object, you would have to overload it: either lose the ability to retry (each run is terminal) or lose idempotency (nothing distinguishes 'same logical run' from 'new run'). The two-object model is what lets restart and once-only completion coexist — which is the whole point of a batch framework.
- How does the choice of identifying parameters directly define a job's idempotency boundary, and what failure mode does getting it wrong cause?Identifying parameters are the instance key, so they decide what counts as 'the same run'. Too coarse (e.g. only the date) and legitimately-different runs collide and hit AlreadyComplete; too fine (e.g. a timestamp) and every launch is a new instance, silently defeating restart and allowing duplicate work.
- Why must the JobRepository sit on a durable, transactional datastore in production for these guarantees to hold?Idempotency and restart depend on persisted instance/execution/context state written transactionally with the work. An in-memory or non-transactional store loses the identity and ExecutionContext on crash, so completed-run detection and resume both break.
saying these in an interview costs you the question
- Treating the split as arbitrary rather than enabling idempotency + restart
- Ignoring that identifying-parameter choice defines the idempotency boundary
- Assuming an in-memory JobRepository is fine for production once-only guarantees
- Thinking the domain model alone guarantees resume without ItemStream-cooperating components