skip to content

What is the JobRepository in Spring Batch, and what does it store?

level: juniorimportance: must knowfreq 70%

answer

  1. Persists batch metadata, not business data
  2. JOB_INSTANCE / JOB_EXECUTION / STEP_EXECUTION / EXECUTION_CONTEXT
  3. Instance = name + identifying params
  4. ExecutionContext enables restart
  5. Needs DataSource + TransactionManager

basics

~10 s

The JobRepository is the component that saves batch execution state to a database. It records which jobs and steps ran, their status, timestamps, and parameters, so a job can be tracked and restarted.

solid answer

~30 s

JobRepository is Spring Batch's central persistence interface for metadata. Every time a job runs, it writes to a set of relational tables: BATCH_JOB_INSTANCE (a logical job identified by name plus identifying parameters), BATCH_JOB_EXECUTION (one row per run attempt with status, start/end time, exit code), BATCH_STEP_EXECUTION (per-step counts and status), and BATCH_JOB/STEP_EXECUTION_CONTEXT (serialized key/value state used for restart). The framework reads this on launch to detect duplicate runs and to resume a failed job from where it stopped. It is backed by a DataSource and a PlatformTransactionManager; an in-memory (map-based) variant historically existed for tests but the JDBC implementation is the norm.

code

java · 17 lines
java
// Spring Batch 5: the repository is auto-provided when you have a DataSource + tx manager.
// Injecting it (or JobExplorer) to inspect metadata:
@Component
class JobStatusReporter {
    private final JobExplorer jobExplorer; // reads the same BATCH_* tables

    JobStatusReporter(JobExplorer jobExplorer) {
        this.jobExplorer = jobExplorer;
    }

    void printLast(String jobName) {
        jobExplorer.getJobInstances(jobName, 0, 1).forEach(instance ->
            jobExplorer.getJobExecutions(instance).forEach(exec ->
                System.out.printf("exec %d status=%s exit=%s%n",
                    exec.getId(), exec.getStatus(), exec.getExitStatus().getExitCode())));
    }
}

go deeper

for a junior

Know it persists run state to BATCH_* tables and enables tracking/restart.

for a middle

Explain instance-vs-execution identity and name the four core tables and their roles.

for a senior

Discuss ExecutionContext serialization, shared DataSource/transaction, and duplicate prevention semantics.

for a principal

Reason about metadata store as a coordination point across nodes and its transactional coupling to business data.

## What it is The `JobRepository` is a Spring Batch interface (`org.springframework.batch.core.repository.JobRepository`) whose job is to **persist and retrieve the metadata of batch executions**. "Metadata" here means the bookkeeping about runs — not the business data you process, but the record of *which* jobs/steps ran, *when*, with *what parameters*, their *status*, and enough saved state to *restart* them. The standard implementation is `SimpleJobRepository`, backed by a set of DAOs that read/write JDBC. It requires a `DataSource` and a `PlatformTransactionManager`. ## The core tables Spring Batch writes to six main tables (all prefixed `BATCH_`): - **BATCH_JOB_INSTANCE** — a *logical* job run. Its identity is the **job name + a hash (JOB_KEY) of the identifying job parameters**. Running the same job with the same identifying parameters refers to the *same* instance. - **BATCH_JOB_EXECUTION** — one row per **attempt** to run an instance. Holds `STATUS` (e.g. `COMPLETED`, `FAILED`), `START_TIME`, `END_TIME`, `EXIT_CODE`, `EXIT_MESSAGE`. A single instance can have many executions (one per restart attempt). - **BATCH_JOB_EXECUTION_PARAMS** — the parameters passed to that execution. - **BATCH_STEP_EXECUTION** — one row per step per job execution, with read/write/commit/skip counts and status. - **BATCH_JOB_EXECUTION_CONTEXT** and **BATCH_STEP_EXECUTION_CONTEXT** — the serialized `ExecutionContext` (a key/value map) at job and step level. This is the durable state that makes **restart** possible (e.g. a reader saving its cursor position or line count). There are also sequence tables/objects (`BATCH_JOB_SEQ`, `BATCH_JOB_EXECUTION_SEQ`, `BATCH_STEP_EXECUTION_SEQ`) for primary keys. ## Why it matters 1. **Restartability** — on failure, the saved `ExecutionContext` lets a step resume rather than reprocess everything. 2. **Duplicate prevention** — because instance identity is name + identifying parameters, the repository can refuse to re-run an already-completed instance (`JobInstanceAlreadyCompleteException`). 3. **Auditing/monitoring** — `JobExplorer`/`JobOperator` read these same tables to report on history. ## Gotchas - Business data and metadata should share the **same transaction/DataSource** where possible so that a chunk commit and its metadata update are atomic. - The metadata schema must exist before the first run (created via the bundled `schema-*.sql` DDL or Boot's `spring.batch.jdbc.initialize-schema`). - In Spring Batch 5 the `ExecutionContext` is serialized as JSON via Jackson by default (older versions used Java serialization).

  • What is the difference between a JobInstance and a JobExecution?
    A JobInstance is a logical job identified by name plus identifying parameters; a JobExecution is a single attempt to run that instance. One instance can have many executions (e.g., a failed attempt followed by a successful restart).
  • Which table makes restart possible?
    BATCH_JOB_EXECUTION_CONTEXT and BATCH_STEP_EXECUTION_CONTEXT — they store the serialized ExecutionContext (saved cursor/counters) so a step can resume where it left off.

saying these in an interview costs you the question

  • Thinking the JobRepository stores the actual business/processed data rather than run metadata
  • Claiming a JobInstance and a JobExecution are the same thing
  • Believing the metadata schema is created automatically with no configuration at all

context