skip to content

Batch Domain Model

The vocabulary the framework persists: Job and Step, JobInstance versus JobExecution, step metrics, the execution context, and the two status types. Restartability only makes sense once you know what is stored, so interviews start here.

part ofSpring Frameworkoverview, primer and where to startread it →
on this pageshow

explore

questions

25

What is the ExecutionContext in Spring Batch, and what is it used for?

level: juniorimportance: must knowfreq 70%

answer

  1. key/value map persisted in JobRepository
  2. step-scoped + job-scoped, not shared
  3. putLong/getString typed accessors
  4. survives restart -> readers resume
  5. small state only

basics

~20 s

ExecutionContext is a persisted key/value map Spring Batch attaches to a running job or step. It stores state — like how many items were read — so if the job stops and restarts, it can pick up where it left off.

solid answer

~40 s

The ExecutionContext is a serializable String-keyed map that Spring Batch persists in the JobRepository. Every StepExecution has its own step-scoped context and every JobExecution has a job-scoped context. Framework components and your code stash small pieces of state in it — item-read counts, file cursor positions, running totals. Because it is written to the database at chunk-commit boundaries, it survives a crash: when you restart a failed job the same context is reloaded, so a reader can skip already-processed items instead of starting over. You read/write it with typed accessors like putLong, putString, put and getLong/getString. It is meant for small state, not for passing large data between steps.

code

java · 11 lines
java
public class CountingTasklet implements Tasklet {
    @Override
    public RepeatStatus execute(StepContribution contribution, ChunkContext chunkContext) {
        ExecutionContext ctx = chunkContext.getStepContext()
                .getStepExecution().getExecutionContext();
        long processed = ctx.getLong("processed", 0L); // second arg = default
        processed += doWork();
        ctx.putLong("processed", processed); // marks context dirty; persisted at commit
        return RepeatStatus.FINISHED;
    }
}

go deeper

for a junior

Should know it is a persisted key/value store used for resuming.

for a middle

Should distinguish step vs job scope and know it lives in the JobRepository DB.

for a senior

Should tie persistence timing (chunk commit) to restart semantics.

for a principal

Should reason about serialization, size limits, and it as a design boundary (not a data bus).

## What it is The `ExecutionContext` (`org.springframework.batch.item.ExecutionContext`) is a **persisted, serializable key/value map** — essentially a `Map<String, Object>` with typed accessors — that Spring Batch uses to carry state across the lifecycle of a batch execution, including across JVM restarts. ## Two scopes There are **two distinct contexts**, and they are not automatically shared: - **Step ExecutionContext** — one per `StepExecution`. Scoped to a single step run. This is where readers store their progress. - **Job ExecutionContext** — one per `JobExecution`. Scoped to the whole job run, useful for sharing a value across steps (with promotion — see the listener question). You obtain them from the `StepExecution`/`JobExecution` objects: `stepExecution.getExecutionContext()` and `stepExecution.getJobExecution().getExecutionContext()`. ## Typed accessors Instead of casting, you use `putLong`, `putInt`, `putDouble`, `putString`, and generic `put(key, obj)`, plus `getLong`, `getString`, `get`, `containsKey`. There is also a `dirty` flag: mutations mark the context dirty so the framework knows it must be re-persisted. ## Where it lives The `JobRepository` serializes the context into database tables — `BATCH_STEP_EXECUTION_CONTEXT` and `BATCH_JOB_EXECUTION_CONTEXT` (columns like `SHORT_CONTEXT` and `SERIALIZED_CONTEXT`). Serialization is handled by an `ExecutionContextSerializer` (Jackson-based by default in modern Spring Batch). ## Why it matters — restartability The step context is persisted **at each chunk commit**. So after processing chunk N, the count/cursor is durable. If the JVM dies mid-job, restarting the same `JobInstance` reloads the saved context and stateful readers (`FlatFileItemReader`, `JdbcCursorItemReader`, etc.) resume from where they stopped rather than reprocessing everything. ## Gotchas - Keep it **small** — the DB columns are size-limited; don't stuff large objects or collections. - Values must be **serializable** by the configured serializer. - Step and job contexts are **separate** — writing to the step context does not make a value visible to later steps. ## When to use Use it for small resumable state and cross-step signals. Do **not** use it as a general data bus for large payloads — that is what item readers/writers and staging tables are for.

  • Where is the ExecutionContext actually stored?
    In the JobRepository — serialized into the BATCH_STEP_EXECUTION_CONTEXT and BATCH_JOB_EXECUTION_CONTEXT tables, via an ExecutionContextSerializer (Jackson-based by default).
  • If I put a value in the step context, can the next step read it?
    No. Step and job contexts are separate. You must promote the key to the job context (e.g. with ExecutionContextPromotionListener) for a later step to see it.

saying these in an interview costs you the question

  • Thinking there is one shared context for the whole job
  • Believing it lives only in memory and is lost on restart
  • Using it to pass large datasets between steps

context

open as a page

In Spring Batch, what is the difference between a JobInstance and a JobExecution?

level: juniorimportance: must knowfreq 78%

basics

~20 s

A JobInstance is a logical run of a job (job name + its identifying parameters). A JobExecution is a single attempt to run that instance. One JobInstance can have several JobExecutions if earlier attempts failed and were restarted.

open as a page

In Spring Batch, what are the Job and Step abstractions, and how do they relate?

level: juniorimportance: must knowfreq 80%

basics

~20 s

A Job is an entire batch process. A Step is one independent phase of it. A Job contains an ordered list of Steps and runs them in sequence; each Step does a unit of work.

open as a page

In Spring Batch, what is the difference between BatchStatus and ExitStatus?

level: juniorimportance: must knowfreq 55%

basics

~20 s

BatchStatus is an enum (COMPLETED, FAILED, STARTED, STOPPED...) describing a job's or step's lifecycle state. ExitStatus is a string-code object (exitCode plus description) used mainly to decide flow between steps. Both live on every JobExecution and StepExecution.

open as a page

What is a StepExecution in Spring Batch, and which metrics/counts does it record?

level: juniorimportance: must knowfreq 55%

basics

~10 s

A StepExecution represents one attempt to run a Step. It records runtime metrics: readCount, writeCount, filterCount, commitCount, rollbackCount, and skipCount, plus status, timestamps, and its own ExecutionContext.

open as a page

Explain the difference between the step-scoped and job-scoped ExecutionContext, and when each is persisted.

level: middleimportance: must knowfreq 60%

basics

~20 s

Each StepExecution has its own step context; each JobExecution has a job context. They are separate. The step context is saved to the database at every chunk commit; the job context is saved as the job execution is updated.

open as a page

What determines whether launching a job creates a new JobInstance or reuses an existing one?

level: middleimportance: must knowfreq 66%

basics

~20 s

The job name plus its identifying JobParameters. If that combination has been seen before, Spring Batch reuses the existing JobInstance; if it is new, it creates a new one. Same job + same identifying params = same instance.

open as a page

How do you construct a Job and a Step using the JobBuilder/StepBuilder DSL in Spring Batch 5?

level: middleimportance: must knowfreq 78%

basics

~20 s

In Spring Batch 5 you instantiate builders directly: new JobBuilder(name, jobRepository) and new StepBuilder(name, jobRepository). The Step builder then picks chunk or tasklet (passing a transaction manager), and the Job builder chains steps with start()/next().

open as a page

How is a step's ExitStatus derived from its BatchStatus by default, and how do you customize ExitStatus?

level: middleimportance: must knowfreq 45%

basics

~10 s

By default the ExitStatus mirrors the BatchStatus (COMPLETED->'COMPLETED', FAILED->'FAILED'). To customize, return a different ExitStatus from a StepExecutionListener's afterStep method (or set it on the StepExecution). The BatchStatus itself stays unchanged.

open as a page

What is StepContribution and how do its counts get aggregated into the StepExecution?

level: seniorimportance: must knowfreq 40%

basics

~10 s

StepContribution is a per-chunk buffer that accumulates read/write/filter/skip counts while a chunk is processed. When the chunk commits, StepExecution.apply(contribution) merges those deltas into the step's running totals.

open as a page

Where do StepExecution and the ExecutionContext live relative to JobInstance and JobExecution, and why does it matter?

level: middleimportance: should knowfreq 40%

basics

~20 s

StepExecutions and the ExecutionContext belong to a JobExecution, not directly to the JobInstance. Each run attempt has its own StepExecutions and context. Because there can be many JobExecutions per instance, each attempt gets a fresh set.

open as a page

What is the difference between filterCount and skipCount on a StepExecution?

level: middleimportance: should knowfreq 45%

basics

~10 s

filterCount counts items an ItemProcessor intentionally dropped by returning null — normal, no error. skipCount counts items that threw an exception and were skipped by a skip policy instead of failing the step.

open as a page

What is ExecutionContextPromotionListener and how do you use it to share data between steps?

level: seniorimportance: should knowfreq 45%

basics

~20 s

It's a step listener that copies chosen keys from a step's ExecutionContext up to the job's ExecutionContext after the step finishes. That's how a value produced in one step becomes readable by a later step.

open as a page

How does the ExecutionContext enable a reader to resume mid-step after a failure? Walk through the ItemStream mechanics.

level: seniorimportance: should knowfreq 50%

basics

~20 s

Stateful readers implement ItemStream. On open() they read a saved position from the ExecutionContext; on update() (called per chunk) they write the current read count. After a crash, restart calls open() with the saved context so the reader skips already-read items.

open as a page

Walk through what happens to JobInstance and JobExecution across a failure and a restart, including the metadata.

level: seniorimportance: should knowfreq 52%

basics

~20 s

The first launch creates one JobInstance and a JobExecution that fails. On restart with the same identifying parameters, Spring Batch reuses the JobInstance, creates a second JobExecution, and resumes using the persisted ExecutionContext. Now the instance has two executions.

open as a page

Explain the AbstractStep template and the Step execution lifecycle. What does AbstractStep handle versus its subclasses?

level: seniorimportance: should knowfreq 45%

basics

~20 s

AbstractStep implements the Step interface and provides a template execute() that handles the surrounding lifecycle — starting/updating the StepExecution, listeners, ExecutionContext, status, and repository saves — while delegating the actual work to an abstract doExecute() implemented by subclasses like TaskletStep.

open as a page

What is the role of AbstractJob and SimpleJob, and how does a SimpleJob execute its ordered list of Steps?

level: seniorimportance: should knowfreq 55%

basics

~20 s

AbstractJob is the base Job implementation holding shared logic (listeners, status, restart). SimpleJob extends it with a plain ordered list of Steps, executing them one at a time in order and stopping if one fails.

open as a page

How does BatchStatus severity ordering work, and what do max(), isRunning(), and isUnsuccessful() do?

level: seniorimportance: should knowfreq 30%

basics

~20 s

BatchStatus values are ordered by severity. BatchStatus.max(a, b) returns the more severe of two statuses — used to aggregate a job's status from its steps. isRunning() is true for STARTING/STARTED; isUnsuccessful() is true for FAILED (and more severe).

open as a page

Where are BatchStatus and ExitStatus persisted, and what practical problems does storing them separately solve?

level: seniorimportance: should knowfreq 22%

basics

~20 s

The JobRepository stores both in the batch metadata tables: a STATUS column holds the BatchStatus name, and an EXIT_CODE/EXIT_MESSAGE column holds the ExitStatus. Keeping them separate lets you record the canonical lifecycle and a flexible outcome code independently.

open as a page

How do commitCount and rollbackCount relate to chunks and transactions, and what makes rollbackCount rise beyond a single failure?

level: seniorimportance: should knowfreq 32%

basics

~20 s

Each chunk runs in one transaction. commitCount increments per successful chunk commit (roughly items/chunkSize). rollbackCount increments each time a chunk transaction rolls back — including the extra rollbacks caused by retry and by single-item re-scanning during skips.

open as a page

Treating a Job as an ordered container of Steps, how do JobInstance/JobExecution/StepExecution identity and the JobRepository enable restart, and what are the design implications of decomposing a process into Steps?

level: principalimportance: should knowfreq 38%

basics

~20 s

A JobInstance is identified by job name + identifying parameters; each attempt is a JobExecution, each step run a StepExecution, all persisted in the JobRepository. On restart of the same instance, completed Steps are skipped so the job resumes mid-way — which is why splitting a process into Steps gives restart granularity.

open as a page

What are the serialization, sizing, and consistency pitfalls of the ExecutionContext, and how do they influence how you use it?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

The context is serialized into size-limited DB columns, so keep it small and serializable. It's saved per chunk with the transaction, so it stays consistent with the data. Don't store large objects or non-serializable types, and namespace keys to avoid collisions.

open as a page

Why did Spring Batch split the run concept into JobInstance and JobExecution instead of a single object, and what does that buy an operator?

level: principalimportance: nice to knowfreq 30%

basics

~10 s

Separating 'the logical run' (JobInstance) from 'each attempt' (JobExecution) lets Spring Batch guarantee a run happens once while still allowing many retry attempts. It enables restartability, idempotency, per-attempt metrics, and a clean audit history.

open as a page

Why does Spring Batch model outcome as two separate concepts (a closed BatchStatus enum and an open ExitStatus code) instead of one?

level: principalimportance: nice to knowfreq 15%

basics

~20 s

Because two different consumers need different vocabularies: the framework needs a small, fixed, trustworthy lifecycle enum for restart and monitoring, while application flows need an open, extensible outcome code for branching and reporting. One field can't safely serve both.

open as a page

How are StepExecution counts aggregated safely across concurrency models — multi-threaded steps versus partitioned steps?

level: principalimportance: nice to knowfreq 20%

basics

~20 s

In a multi-threaded step, many chunks share one StepExecution and merge counts via the synchronized apply(StepContribution). In partitioning, each partition is its own child StepExecution with independent counts; the manager step sums them for reporting.

open as a page