Explain the AbstractStep template and the Step execution lifecycle. What does AbstractStep handle versus its subclasses?
answer
- AbstractStep = template execute(); subclass doExecute()
- Lifecycle: status STARTED -> open ExecutionContext -> beforeStep -> doExecute -> status/ExitStatus -> afterStep -> repo update
- TaskletStep is the main subclass (repeat + tx loop)
- ExecutionContext restored on restart
- startLimit + allowStartIfComplete govern restart
basics
~20 sAbstractStep implements the Step interface and provides a template execute() that handles the surrounding lifecycle — starting/updating the StepExecution, listeners, ExecutionContext, status, and repository saves — while delegating the actual work to an abstract doExecute() implemented by subclasses like TaskletStep.
solid answer
~40 s`AbstractStep` is the abstract base implementing the `Step` interface, using the Template Method pattern. Its `execute(StepExecution)` owns the invariant step lifecycle: set status STARTED, persist the StepExecution, open the step-scoped `ExecutionContext` (restoring saved state on restart), fire `StepExecutionListener.beforeStep`, then call the abstract **`doExecute(StepExecution)`** for the real work, then compute BatchStatus/ExitStatus, fire `afterStep`, close resources, and update the JobRepository — with exception handling that marks the step FAILED and still persists it. Subclasses provide `doExecute`: the main one is **`TaskletStep`**, whose doExecute drives the repeat/transaction loop calling the `Tasklet` (chunk processing is a chunk-oriented tasklet). Other subclasses include `PartitionStep`, `FlowStep`, and `JobStep`. So AbstractStep guarantees consistent metadata, listener callbacks, and restart-state handling regardless of step type.
code
java · 13 lines// A Tasklet is what TaskletStep.doExecute() drives inside its repeat/transaction loop.
@Bean
public Step archiveStep(JobRepository jobRepository, PlatformTransactionManager txManager) {
return new StepBuilder("archiveStep", jobRepository)
.tasklet((contribution, chunkContext) -> {
// real work goes here; AbstractStep wraps status/listeners/context around it
archiveYesterdaysFiles();
return RepeatStatus.FINISHED; // tells the repeat loop we're done
}, txManager)
.startLimit(3) // Step.getStartLimit(): max starts across restarts
.allowStartIfComplete(false) // don't re-run if already COMPLETED
.build(); // -> a TaskletStep (extends AbstractStep)
}go deeper
Know a Step has a lifecycle and that AbstractStep/TaskletStep are the base classes behind the builder.
Describe the execute() lifecycle at a high level and that TaskletStep does the actual repeat/transaction work.
Explain the template-method split, ExecutionContext restore on restart, listener timing, and startLimit/allowStartIfComplete.
Reason about restart correctness (ItemStream state), duplicate-work risks from allowStartIfComplete, and choosing TaskletStep vs Partition/Flow/Job steps.
**`Step` is an interface**; **`AbstractStep`** (`org.springframework.batch.core.step.AbstractStep`) is the abstract base class that virtually all step types extend. It applies the **Template Method** pattern so every step gets identical lifecycle bookkeeping while differing only in the actual work performed. **The Step interface** exposes: `getName()`, `execute(StepExecution stepExecution)`, `getStartLimit()` (how many times the step may be started across restarts), and `isAllowStartIfComplete()` (whether to re-run even if previously COMPLETED). **What AbstractStep.execute() does (the invariant lifecycle):** 1. Sets the `StepExecution` `BatchStatus` to STARTED and its start time, then persists it via the `JobRepository`. 2. Opens the step-scoped **`ExecutionContext`** through an `ItemStream`/context mechanism — on a restart this **restores previously saved state** (e.g., the last read position), which is central to resuming work. 3. Fires `StepExecutionListener.beforeStep(stepExecution)`. 4. Calls the abstract **`doExecute(StepExecution)`** — the subclass hook where real processing happens. 5. On normal completion or exception, computes the final `BatchStatus` and `ExitStatus`. Exceptions are caught, logged, the step marked FAILED, and the exception recorded — the step is still persisted so a restart can see it. 6. Fires `StepExecutionListener.afterStep(...)` (which may adjust the ExitStatus). 7. Closes streams/resources and does a final `JobRepository.update(stepExecution)`. Because all of this lives in the base, **every** step type gets consistent metadata, listener callbacks, and restart semantics for free. **What subclasses provide (`doExecute`):** - **`TaskletStep`** — the workhorse. Its `doExecute` runs a `RepeatTemplate` loop; each iteration opens a transaction (via the `PlatformTransactionManager`), invokes the `Tasklet.execute(...)`, and commits — repeating until the tasklet returns `RepeatStatus.FINISHED`. Chunk-oriented processing is implemented as a `ChunkOrientedTasklet` running inside a TaskletStep, which is why chunk internals belong to a sibling topic while the *container* is TaskletStep. - **`PartitionStep`** — splits work into partitions executed (possibly in parallel) by a `PartitionHandler`. - **`FlowStep`** — wraps a `Flow` as a single step. - **`JobStep`** — runs an entire nested `Job` as one step. **When to care:** you rarely subclass AbstractStep yourself; you configure a step via `StepBuilder`, which assembles a `TaskletStep`. Understanding the template explains *why* listeners fire when they do, *why* the ExecutionContext is restored on restart, and *why* even a crashing step leaves a persisted FAILED StepExecution enabling restart. **Gotchas:** 1. `getStartLimit()` (default from the builder) caps how many times a step may be *started*; exceeding it on repeated restarts throws `StartLimitExceededException` — a step that keeps failing won't retry forever. 2. `allowStartIfComplete(true)` makes a completed step re-run on every restart — useful for setup/validation steps, dangerous for data-loading steps (duplicate work). 3. The ExecutionContext restore only helps if your reader/writer participate as `ItemStream`s and actually persist their position; a custom component that ignores the ExecutionContext will not resume correctly. 4. Listener `afterStep` runs even when the step failed (so cleanup/alerting can happen), and it can override the ExitStatus — a subtle source of 'why did my exit code change' confusion.
- Where does chunk-oriented (reader/processor/writer) processing fit relative to AbstractStep and TaskletStep?Chunk processing is a ChunkOrientedTasklet running inside a TaskletStep (which extends AbstractStep). So the step container is TaskletStep; the chunk read-process-write loop and commit-interval internals are the tasklet's job — a separate concern from the step abstraction itself.
- What does getStartLimit() protect against?It caps how many times a step may be started across restarts; once exceeded, restarting throws StartLimitExceededException, preventing an endlessly failing step from being retried indefinitely.
saying these in an interview costs you the question
- Saying AbstractStep contains the reader/processor/writer loop (that's the tasklet/chunk layer)
- Claiming listeners' afterStep does not run when the step fails
- Thinking a failed step leaves no persisted StepExecution
- Believing allowStartIfComplete has no effect on restart behavior