skip to content

JobRepository & Metadata Schema

The JobRepository persists instances, executions and contexts to a documented set of tables, which is what makes restart and monitoring possible. Interviewers ask where that state lives and what happens if you point it at an in-memory database.

part ofSpring Frameworkoverview, primer and where to startread it →
on this pageshow

explore

questions

5

What is the JobRepository in Spring Batch, and what does it store?

level: juniorimportance: must knowfreq 70%

answer

  1. Persists batch metadata, not business data
  2. JOB_INSTANCE / JOB_EXECUTION / STEP_EXECUTION / EXECUTION_CONTEXT
  3. Instance = name + identifying params
  4. ExecutionContext enables restart
  5. Needs DataSource + TransactionManager

basics

~10 s

The JobRepository is the component that saves batch execution state to a database. It records which jobs and steps ran, their status, timestamps, and parameters, so a job can be tracked and restarted.

solid answer

~30 s

JobRepository is Spring Batch's central persistence interface for metadata. Every time a job runs, it writes to a set of relational tables: BATCH_JOB_INSTANCE (a logical job identified by name plus identifying parameters), BATCH_JOB_EXECUTION (one row per run attempt with status, start/end time, exit code), BATCH_STEP_EXECUTION (per-step counts and status), and BATCH_JOB/STEP_EXECUTION_CONTEXT (serialized key/value state used for restart). The framework reads this on launch to detect duplicate runs and to resume a failed job from where it stopped. It is backed by a DataSource and a PlatformTransactionManager; an in-memory (map-based) variant historically existed for tests but the JDBC implementation is the norm.

code

java · 17 lines
java
// Spring Batch 5: the repository is auto-provided when you have a DataSource + tx manager.
// Injecting it (or JobExplorer) to inspect metadata:
@Component
class JobStatusReporter {
    private final JobExplorer jobExplorer; // reads the same BATCH_* tables

    JobStatusReporter(JobExplorer jobExplorer) {
        this.jobExplorer = jobExplorer;
    }

    void printLast(String jobName) {
        jobExplorer.getJobInstances(jobName, 0, 1).forEach(instance ->
            jobExplorer.getJobExecutions(instance).forEach(exec ->
                System.out.printf("exec %d status=%s exit=%s%n",
                    exec.getId(), exec.getStatus(), exec.getExitStatus().getExitCode())));
    }
}

go deeper

for a junior

Know it persists run state to BATCH_* tables and enables tracking/restart.

for a middle

Explain instance-vs-execution identity and name the four core tables and their roles.

for a senior

Discuss ExecutionContext serialization, shared DataSource/transaction, and duplicate prevention semantics.

for a principal

Reason about metadata store as a coordination point across nodes and its transactional coupling to business data.

## What it is The `JobRepository` is a Spring Batch interface (`org.springframework.batch.core.repository.JobRepository`) whose job is to **persist and retrieve the metadata of batch executions**. "Metadata" here means the bookkeeping about runs — not the business data you process, but the record of *which* jobs/steps ran, *when*, with *what parameters*, their *status*, and enough saved state to *restart* them. The standard implementation is `SimpleJobRepository`, backed by a set of DAOs that read/write JDBC. It requires a `DataSource` and a `PlatformTransactionManager`. ## The core tables Spring Batch writes to six main tables (all prefixed `BATCH_`): - **BATCH_JOB_INSTANCE** — a *logical* job run. Its identity is the **job name + a hash (JOB_KEY) of the identifying job parameters**. Running the same job with the same identifying parameters refers to the *same* instance. - **BATCH_JOB_EXECUTION** — one row per **attempt** to run an instance. Holds `STATUS` (e.g. `COMPLETED`, `FAILED`), `START_TIME`, `END_TIME`, `EXIT_CODE`, `EXIT_MESSAGE`. A single instance can have many executions (one per restart attempt). - **BATCH_JOB_EXECUTION_PARAMS** — the parameters passed to that execution. - **BATCH_STEP_EXECUTION** — one row per step per job execution, with read/write/commit/skip counts and status. - **BATCH_JOB_EXECUTION_CONTEXT** and **BATCH_STEP_EXECUTION_CONTEXT** — the serialized `ExecutionContext` (a key/value map) at job and step level. This is the durable state that makes **restart** possible (e.g. a reader saving its cursor position or line count). There are also sequence tables/objects (`BATCH_JOB_SEQ`, `BATCH_JOB_EXECUTION_SEQ`, `BATCH_STEP_EXECUTION_SEQ`) for primary keys. ## Why it matters 1. **Restartability** — on failure, the saved `ExecutionContext` lets a step resume rather than reprocess everything. 2. **Duplicate prevention** — because instance identity is name + identifying parameters, the repository can refuse to re-run an already-completed instance (`JobInstanceAlreadyCompleteException`). 3. **Auditing/monitoring** — `JobExplorer`/`JobOperator` read these same tables to report on history. ## Gotchas - Business data and metadata should share the **same transaction/DataSource** where possible so that a chunk commit and its metadata update are atomic. - The metadata schema must exist before the first run (created via the bundled `schema-*.sql` DDL or Boot's `spring.batch.jdbc.initialize-schema`). - In Spring Batch 5 the `ExecutionContext` is serialized as JSON via Jackson by default (older versions used Java serialization).

  • What is the difference between a JobInstance and a JobExecution?
    A JobInstance is a logical job identified by name plus identifying parameters; a JobExecution is a single attempt to run that instance. One instance can have many executions (e.g., a failed attempt followed by a successful restart).
  • Which table makes restart possible?
    BATCH_JOB_EXECUTION_CONTEXT and BATCH_STEP_EXECUTION_CONTEXT — they store the serialized ExecutionContext (saved cursor/counters) so a step can resume where it left off.

saying these in an interview costs you the question

  • Thinking the JobRepository stores the actual business/processed data rather than run metadata
  • Claiming a JobInstance and a JobExecution are the same thing
  • Believing the metadata schema is created automatically with no configuration at all

context

open as a page

How does the Spring Batch metadata schema get created, and what is the role of the schema-*.sql files and @EnableBatchProcessing?

level: middleimportance: must knowfreq 60%

basics

~10 s

Spring Batch ships DDL scripts named schema-<database>.sql inside its jar. Spring Boot runs the right one automatically based on spring.batch.jdbc.initialize-schema. @EnableBatchProcessing wires up the batch infrastructure (JobRepository, etc.).

open as a page

Walk through the relationships between BATCH_JOB_INSTANCE, JOB_EXECUTION, STEP_EXECUTION, and EXECUTION_CONTEXT, and how they enable restart.

level: seniorimportance: should knowfreq 45%

basics

~20 s

One JobInstance (name + identifying params) has many JobExecutions (one per attempt). Each JobExecution has many StepExecutions. ExecutionContext rows store saved state per job and per step. On restart, Spring reuses the instance and reloads the ExecutionContext to resume.

open as a page

Why does the JobRepository use a SERIALIZABLE isolation level when creating a JobExecution, and when would you change it?

level: seniorimportance: should knowfreq 40%

basics

~20 s

When starting a job, the repository first checks whether an instance/execution already exists, then inserts one. It uses SERIALIZABLE isolation so two launchers can't both create the same JobInstance at once. You lower it only if your DB struggles with SERIALIZABLE.

open as a page

What production concerns arise from the JobRepository metadata store, and how do you configure it for a clustered deployment?

level: principalimportance: nice to knowfreq 25%

basics

~20 s

Use a shared, persistent database (not an in-memory store) so all nodes see the same metadata. Manage the schema via Flyway/Liquibase (initialize-schema=never), share one DataSource/transaction manager, and rely on isolation-on-create plus locks to prevent duplicate concurrent launches.

open as a page