skip to content

Restartability

A restart resumes a failed or stopped execution from the last committed chunk using the persisted execution context, subject to restartable, startLimit and allowStartIfComplete flags. Interviewers ask what makes a job safely restartable, and readers with saved positions are the core of it.

part ofSpring Frameworkoverview, primer and where to startread it →
on this pageshow

explore

questions

5

What does it mean to restart a Spring Batch job, and which jobs can be restarted?

level: juniorimportance: must knowfreq 60%

answer

  1. FAILED/STOPPED → new JobExecution, same JobInstance
  2. COMPLETED → JobInstanceAlreadyCompleteException
  3. JobInstance = name + identifying params
  4. unique param each run = never restarts
  5. needs persistent JobRepository

basics

~20 s

Restarting means re-running a job that previously FAILED or STOPPED. Spring Batch reuses the same JobInstance (same identifying parameters) and continues where it left off instead of starting from scratch. A COMPLETED job cannot normally be restarted.

solid answer

~40 s

A Spring Batch job run is a JobExecution belonging to a JobInstance (identified by the job name plus its identifying JobParameters). If a JobExecution ends FAILED or STOPPED, launching the job again with the same identifying parameters creates a NEW JobExecution under the SAME JobInstance — that is a restart. Spring Batch reads the persisted metadata (JobRepository) to know which steps completed and where the failing step stopped, so it resumes rather than redoing finished work. A JobInstance that already COMPLETED cannot be restarted — relaunching with identical identifying parameters throws JobInstanceAlreadyCompleteException. Restart requires the JobRepository to have persisted state, so an in-memory-only setup loses restartability.

code

java · 14 lines
java
// Same identifying parameters on both attempts => same JobInstance => restart
JobParameters params = new JobParametersBuilder()
        .addString("inputFile", "orders-2026-07-22.csv") // identifying
        .toJobParameters();

// First launch: FAILED at some chunk
JobExecution first = jobLauncher.run(importJob, params);

// Later, relaunch with the SAME params => new JobExecution, same JobInstance,
// resumes instead of starting over.
JobExecution restart = jobLauncher.run(importJob, params);

// If 'first' had COMPLETED, this second call throws
// JobInstanceAlreadyCompleteException instead.

go deeper

for a junior

Know: restart = re-run a FAILED/STOPPED job with the same parameters; it continues rather than redoing everything.

for a middle

Explain JobInstance vs JobExecution and why identifying parameters must stay stable across attempts.

for a senior

Discuss the JobRepository persistence requirement and the exact exceptions (JobInstanceAlreadyComplete / AlreadyRunning).

for a principal

Reason about JobInstance identity design, idempotency vs restart trade-offs, and metadata durability across crashes.

## Core vocabulary - **Job**: the definition of a batch process (a bean of type `Job`). - **JobParameters**: key/value inputs to a run. Parameters can be *identifying* (default) or *non-identifying*. The *identifying* ones, together with the job name, define the **JobInstance**. - **JobInstance**: a logical *run of a job for a given input*. Example: `importFile` for `date=2026-07-22` is one JobInstance; a different date is a different JobInstance. - **JobExecution**: a single *attempt* to run a JobInstance. One JobInstance can have MANY JobExecutions — one per attempt. Each has a `BatchStatus` (`STARTING`, `STARTED`, `COMPLETED`, `FAILED`, `STOPPED`, `ABANDONED`). - **JobRepository**: the persistence layer (normally the `BATCH_*` metadata tables) that stores JobInstances, JobExecutions, StepExecutions and their ExecutionContexts. ## What restart actually is "Restart" is not a special API call — you simply **launch the same job again with the same identifying JobParameters**. Spring Batch looks up the JobInstance: - If no JobInstance exists → it creates one and runs the first JobExecution. - If a JobInstance exists whose last JobExecution is `FAILED` or `STOPPED` → it creates a **new JobExecution** under that **same** JobInstance. This is the restart. The framework consults persisted metadata to skip already-`COMPLETED` steps and to resume the step that was in progress. - If a JobInstance exists whose last JobExecution is `COMPLETED` → by default you get a `JobInstanceAlreadyCompleteException` ("a completed instance cannot be restarted"). - If a JobExecution is currently `STARTED`/running → you may get a `JobExecutionAlreadyRunningException`. ## Why identifying parameters matter Because the JobInstance is keyed by name + identifying parameters, adding a unique parameter (e.g. `run.id` or a timestamp) each time produces a **brand-new JobInstance every launch** — which means you are always starting fresh and **never restarting**. This is a common way people accidentally disable restart. If you want restartability, keep the identifying parameters stable across attempts. ## Persistence requirement Restart depends entirely on the JobRepository having durable state. With a real database (the normal setup) metadata survives a crash, so you can restart after a JVM death. If the repository is in-memory/transient, a crash loses the metadata and there is nothing to resume from. ## Terminal vs restartable states - Restartable end states: `FAILED`, `STOPPED`. - Not restartable: `COMPLETED` (finished successfully) and `ABANDONED` (explicitly marked as not-to-be-rerun). ## When to use Restart is the backbone of resilient batch: a long import that dies at record 800k should resume near 800k, not reprocess from zero. You get this essentially for free with chunk-oriented steps and a persistent JobRepository, provided you don't sabotage the JobInstance identity.

  • What happens if you add a unique run.id parameter on every launch?
    Each launch has different identifying parameters, so a new JobInstance is created every time. You never restart — every run starts from scratch. This is fine for idempotent jobs but defeats restart.
  • Can you restart a COMPLETED job?
    Not by default. Relaunching the same JobInstance after COMPLETED throws JobInstanceAlreadyCompleteException. You'd need different identifying parameters (a new instance), or per-step allowStartIfComplete to re-run specific steps within a restart.

saying these in an interview costs you the question

  • Thinking restart is a dedicated API/method rather than relaunching with the same parameters
  • Believing a COMPLETED job can be restarted with identical parameters
  • Adding a timestamp/UUID parameter every run and still expecting restart to work
  • Assuming restart works with a purely in-memory JobRepository after a crash

context

open as a page

How does a chunk-oriented step resume at the last good chunk on restart? What role does ExecutionContext play?

level: middleimportance: must knowfreq 55%

basics

~20 s

Each committed chunk saves progress (like the current read position) into the step's persisted ExecutionContext. On restart, the reader is re-opened and reads that saved state, so processing continues from just after the last successfully committed chunk instead of the beginning.

open as a page

What does the job-level restartable flag do, and what happens if you try to restart a non-restartable job?

level: middleimportance: should knowfreq 35%

basics

~10 s

Setting a job's restartable property to false marks it as run-once-per-instance. If its execution fails and you relaunch with the same identifying parameters, Spring Batch refuses and throws JobRestartException instead of resuming.

open as a page

Explain the step-level startLimit and allowStartIfComplete settings and how they interact on restart.

level: seniorimportance: should knowfreq 30%

basics

~20 s

startLimit caps how many times a step may be started across all executions of a JobInstance (default effectively unlimited); exceeding it fails with StartLimitExceededException. allowStartIfComplete=true forces a step that already COMPLETED to run again on every restart instead of being skipped.

open as a page

You inherit a nightly batch job that 'never resumes' after failures — it always reprocesses everything. Walk through the likely causes and how restart correctness depends on job/step design.

level: principalimportance: nice to knowfreq 20%

basics

~20 s

Most often the job adds a unique identifying parameter (timestamp/UUID) each run, so every launch is a NEW JobInstance and there's nothing to resume. Other causes: readers with saveState=false or no ItemStream, an in-memory JobRepository losing state, or restartable=false.

open as a page