skip to content

In Spring Batch, what are the Job and Step abstractions, and how do they relate?

level: juniorimportance: must knowfreq 80%

answer

  1. Job = whole process; Step = one phase
  2. Job is an ordered container of Steps
  3. JobBuilder.start(a).next(b).build()
  4. JobExecution per run, StepExecution per step
  5. Fail stops job; restart skips completed steps

basics

~20 s

A Job is an entire batch process. A Step is one independent phase of it. A Job contains an ordered list of Steps and runs them in sequence; each Step does a unit of work.

solid answer

~40 s

In Spring Batch, `Job` and `Step` are the two core domain abstractions. A `Job` represents the whole batch process end to end (for example 'import daily orders'); it is essentially an ordered container of `Step`s. A `Step` is a self-contained, independent phase of that Job — read/transform/write a chunk of data, or run a single tasklet. The Job orchestrates: it runs its Steps in the order declared, and if a Step fails the Job typically stops with that failure. You define both with builders — `JobBuilder` and `StepBuilder`. Each run of a Job produces a `JobExecution`, and each Step run produces a `StepExecution`, which Spring Batch persists in the JobRepository so a failed Job can be restarted and completed Steps skipped.

code

java · 7 lines
java
@Bean
public Job importJob(JobRepository jobRepository, Step loadStep, Step reportStep) {
    return new JobBuilder("importJob", jobRepository)
            .start(loadStep)   // first step
            .next(reportStep)  // then this one, in order
            .build();          // -> a SimpleJob (a Job)
}

go deeper

for a junior

Know Job = whole process, Step = one phase, and that a Job runs its Steps in order.

for a middle

Add that each run produces a JobExecution/StepExecution persisted in the JobRepository, enabling restart.

for a senior

Discuss restart granularity, why steps are independent, and sequential vs flow orchestration.

for a principal

Frame step decomposition as a design decision balancing restart granularity, failure isolation, and operational visibility.

Spring Batch models a batch process with two central abstractions defined as interfaces: **`Job`** (`org.springframework.batch.core.Job`) and **`Step`** (`org.springframework.batch.core.Step`). ## Job The entire batch process as one logical unit of work, from start to finish. Think 'end-of-day settlement' or 'nightly CSV import'. A `Job` is fundamentally an **ordered container of `Step`s**: it declares which steps run and in what order, and it owns the overall success/failure/restart semantics. The `Job` interface exposes `getName()`, `execute(JobExecution)`, `isRestartable()`, `getJobParametersIncrementer()`, and `getJobParametersValidator()`. ## Step A single, independent, sequential phase of a Job. Each step encapsulates one distinct piece of processing — commonly: - a chunk-oriented step (read → process → write) or - a tasklet step (one arbitrary action like a file move). Steps are meant to be independent so they can succeed, fail, and restart on their own. The `Step` interface exposes `getName()`, `execute(StepExecution)`, `isAllowStartIfComplete()`, and `getStartLimit()`. ## How they relate A Job holds an ordered list of Steps and runs them one after another. When you write `new JobBuilder("import", jobRepository).start(loadStep).next(reportStep).build()`, you get a Job that runs `loadStep`, then `reportStep`. If a step fails, the job stops there and is marked FAILED; a later restart of the same JobInstance skips already-COMPLETED steps and resumes at the failed one. ## Runtime vs definition The Job/Step you build are the *definitions* (blueprints). Each actual run creates runtime metadata: - a `JobInstance` (a logical run identified by job name + identifying JobParameters), - a `JobExecution` (one attempt at that instance), - and a `StepExecution` per step attempt. Spring Batch persists all of this in the **JobRepository**, which is what makes restartability possible. ## When to use which Reach for multiple Steps when your process has genuinely distinct phases (extract, then transform, then notify) — this gives you restart granularity and clearer failure isolation. Keep logic inside a single Step when it is one cohesive read/write flow. **Gotchas:** - (1) Steps in a plain `SimpleJob` are strictly sequential; branching/conditional flow requires the flow DSL (`FlowJob`). - (2) A Job is not a Step and a Step is not a Job — though you can nest a Job inside a Step via `JobStep`, that is a distinct construct, not the default. - (3) The Job doesn't 'do work' itself; all real processing lives in Steps.

  • If the second Step fails and you restart the Job, does the first Step run again?
    No. On restart of the same JobInstance, Steps already marked COMPLETED are skipped, and execution resumes at the failed step — unless that step is configured with allowStartIfComplete=true.
  • Does the Job itself perform the reading and writing?
    No. The Job only orchestrates order and overall status; the actual read/process/write work lives inside the Steps.

saying these in an interview costs you the question

  • Saying a Job and a Step are interchangeable or the same thing
  • Claiming the Job does the data processing itself
  • Thinking Steps in a SimpleJob can run in parallel by default
  • Believing a restart re-runs all steps from scratch

context