skip to content

What production concerns arise from the JobRepository metadata store, and how do you configure it for a clustered deployment?

level: principalimportance: nice to knowfreq 25%

answer

  1. Shared durable RDBMS — no in-memory store in prod
  2. initialize-schema=never + Flyway/Liquibase
  3. Same DataSource/tx manager as business writes → atomic
  4. Clustered launch: serializable-on-create + unique constraint / ShedLock
  5. Retention pruning + JobOperator abandon/restart

basics

~20 s

Use a shared, persistent database (not an in-memory store) so all nodes see the same metadata. Manage the schema via Flyway/Liquibase (initialize-schema=never), share one DataSource/transaction manager, and rely on isolation-on-create plus locks to prevent duplicate concurrent launches.

solid answer

~40 s

In production the JobRepository must be backed by a **shared, durable relational database** so every node/pod observes the same instance/execution state — an in-memory or per-node store breaks restart and duplicate detection. Own the schema through **Flyway/Liquibase** and set `spring.batch.jdbc.initialize-schema=never` so migrations are versioned, not auto-run. The repository's `DataSource` and `PlatformTransactionManager` should ideally be the **same** ones your business writes use, so a chunk's data and its metadata commit atomically. For clustered launches, the create path's `ISOLATION_SERIALIZABLE` prevents two nodes creating the same instance; if you lower it for DB compatibility, add a **unique constraint or distributed lock** (e.g. ShedLock) so only one node launches a given job. Also plan **retention/pruning** of old executions, monitor via `JobExplorer`/`JobOperator`, and consider `ExecutionContext` serialization (JSON/Jackson) size limits.

code

java · 23 lines
java
// Production wiring: explicit shared DataSource + transaction manager, schema owned by Flyway.
@Configuration
@EnableBatchProcessing(
    dataSourceRef = "batchDataSource",
    transactionManagerRef = "batchTxManager",
    tablePrefix = "BATCH_")
class ProdBatchInfra {
    @Bean DataSource batchDataSource(/* pooled, shared with business writes */) { /* ... */ return null; }
    @Bean PlatformTransactionManager batchTxManager(DataSource ds) {
        return new DataSourceTransactionManager(ds);
    }
}

// application.properties
// spring.batch.jdbc.initialize-schema=never   # Flyway/Liquibase owns the schema

// Guard clustered launches with a distributed lock (e.g. ShedLock) so only one node starts a given job.
@Scheduled(cron = "0 0 * * * *")
@SchedulerLock(name = "nightlyImportJob", lockAtMostFor = "30m")
void launch() throws Exception {
    jobLauncher.run(importJob, new JobParametersBuilder()
        .addLong("run.id", System.currentTimeMillis()).toJobParameters());
}

go deeper

for a junior

Know it must be a real shared database in production, not in-memory.

for a middle

Explain schema ownership via migrations and initialize-schema=never.

for a senior

Discuss shared DataSource/tx coupling, concurrent-launch guards, and JobOperator recovery.

for a principal

Design the metadata store as clustered coordination state: isolation vs external locks, transactional atomicity, retention, upgrade migrations, and serializer/pool interactions.

## Why the metadata store is a first-class production concern The `JobRepository` is effectively **shared coordination state**. Restart, duplicate prevention, and monitoring all depend on every process reading a **single, consistent** copy of it. Get this wrong and you get double-processing, lost restarts, or corrupt history. ## 1. Persistent, shared database - Never rely on an ephemeral/in-memory metadata store in production (the old `MapJobRepository` was removed in Batch 5 precisely because it misled people). Use a **real RDBMS** (Postgres, MySQL, Oracle, SQL Server) reachable by all nodes. - All app instances point at the **same** database and table prefix so instance identity is global, not per-node. ## 2. Schema ownership - Set `spring.batch.jdbc.initialize-schema=never` and apply the bundled `schema-<platform>.sql` through **Flyway or Liquibase**. This makes schema changes reviewed, versioned, and repeatable across environments — and avoids surprise DDL on startup. - Keep an eye on **cross-version schema migrations** when upgrading Spring Batch (e.g. Batch 4→5 changed some columns/types); the release notes ship migration DDL. ## 3. Transactional coupling to business data - Prefer the **same `DataSource` + `PlatformTransactionManager`** for metadata and business writes. Then a chunk's item writes and the `StepExecution`/`ExecutionContext` update commit in one transaction — so counts and saved position never diverge from the data actually written. Splitting them across two datasources reintroduces the possibility of "wrote data but lost the metadata" on a crash. ## 4. Concurrent / clustered launches - The create path's default `ISOLATION_SERIALIZABLE` stops two nodes from both creating the same `JobInstance`. - If you must lower it (DB doesn't support serializable well, deadlocks), compensate with: - a **unique constraint** on the instance (name + job key), - a **distributed lock** (e.g. ShedLock, database advisory lock) around the launch, - or a **single-launcher** design (one scheduler leader, or message-driven single consumer). - For scaling *within* a job (remote partitioning / remote chunking), the shared JobRepository is what workers use to persist their `StepExecution`s — another reason it must be shared and correctly isolated. ## 5. Operational lifecycle - **Retention:** these tables grow unbounded. Plan pruning/archival of old `JOB_EXECUTION`/`STEP_EXECUTION`/context rows (respecting FKs; delete children first). Spring Batch doesn't auto-purge. - **Monitoring/recovery:** use `JobExplorer` (read-only) and `JobOperator` (`stop`, `restart`, `abandon`) against the same store. Stuck `STARTED` executions after a crash may need `abandon` before restart. - **ExecutionContext size:** it's stored in a limited-width column (historically ~2500 chars short context / a larger serialized column). Huge contexts cause truncation/serialization errors — keep saved state small. - **Serializer choice:** Batch 5 defaults to `Jackson2ExecutionContextStringSerializer` (JSON); ensure all persisted keys are JSON-serializable and stable across deploys. ## 6. Isolation & connection pool interactions - Serializable-on-create can hold locks; size your pool and set sane lock timeouts so a launch storm doesn't exhaust connections. Some managed DBs need an explicit lower level plus an external lock. ## When to deviate - Truly single-node, non-restartable, throwaway jobs can tolerate simpler setups — but the moment you have multiple replicas, schedulers, or restart requirements, the shared-persistent-store discipline above is mandatory.

  • Why prefer the same DataSource and transaction manager for metadata and business writes?
    So a chunk's item writes and its StepExecution/ExecutionContext update commit in a single transaction. Splitting them across two datasources means a crash can leave metadata and actual data out of sync, breaking counts and restart correctness.
  • After a pod crash, a JobExecution is stuck in STARTED and you can't restart it. What do you do?
    Use JobOperator to abandon (mark ABANDONED) the orphaned STARTED execution, then restart the instance. The stuck row exists because the process died before writing a terminal status.
  • How do you keep the metadata tables from growing without bound?
    Implement retention/pruning — periodically delete old executions (children first to respect FKs) or archive them. Spring Batch has no built-in purge; teams script it or use scheduled cleanup jobs.

saying these in an interview costs you the question

  • Using an in-memory / per-node metadata store in a multi-replica production deployment
  • Leaving initialize-schema=always in production instead of versioned migrations
  • Assuming SERIALIZABLE-on-create is enough when you've lowered it without any other lock
  • Ignoring unbounded growth of the BATCH_ tables
  • Storing large blobs in the ExecutionContext

context