skip to content

Explain the step-level startLimit and allowStartIfComplete settings and how they interact on restart.

level: seniorimportance: should knowfreq 30%

answer

  1. startLimit = max starts per JobInstance → StartLimitExceededException
  2. allowStartIfComplete=true → re-run COMPLETED step each restart
  3. default: completed step skipped on restart
  4. startLimit counts starts not failures
  5. not the same as skipLimit/retryLimit (fault tolerance)

basics

~20 s

startLimit caps how many times a step may be started across all executions of a JobInstance (default effectively unlimited); exceeding it fails with StartLimitExceededException. allowStartIfComplete=true forces a step that already COMPLETED to run again on every restart instead of being skipped.

solid answer

~40 s

Both are per-step controls that shape restart behavior. `startLimit(n)` bounds how many times a given step may be *started* across all restarts of the same JobInstance; each restart that runs the step counts against it, and once exceeded Spring Batch throws `StartLimitExceededException` — useful to stop endlessly retrying a step that keeps failing. `allowStartIfComplete(true)` overrides the default rule that a `COMPLETED` step is skipped on restart: with it, the step re-runs on every restart of the instance. You use allowStartIfComplete for steps that must always execute — validation, setup, cleanup, or non-restartable resources you want reprocessed. They interact: allowStartIfComplete re-runs a completed step, but startLimit still counts each start, so an always-restarting step can eventually hit its limit. Neither overrides the job-level restartable flag.

code

java · 21 lines
java
@Bean
public Job pipelineJob(JobRepository repo, PlatformTransactionManager tx,
                       Tasklet validate, ItemReader<Rec> r, ItemWriter<Rec> w) {
    Step validateStep = new StepBuilder("validate", repo)
            .tasklet(validate, tx)
            .allowStartIfComplete(true)   // always re-validate on restart
            .build();

    Step loadStep = new StepBuilder("load", repo)
            .<Rec, Rec>chunk(500, tx)
            .reader(r).writer(w)
            .startLimit(3)                // at most 3 starts across restarts
            .build();

    return new JobBuilder("pipelineJob", repo)
            .start(validateStep)
            .next(loadStep)
            .build();
    // Restart: validate runs again (allowStartIfComplete); load resumes at
    // its last chunk, but a 4th start of load throws StartLimitExceededException.
}

go deeper

for a junior

Aware that steps can be limited in how often they restart and that some steps can be forced to re-run.

for a middle

Define both properties and their defaults; give a use case for each.

for a senior

Explain the interaction (startLimit counts allowStartIfComplete re-runs), the exceptions, and idempotency implications.

for a principal

Design restart topology: which steps re-run vs resume, bounding environmental retries, and separating restart controls from item-level fault tolerance.

## Default restart behavior for steps When a JobInstance restarts, Spring Batch walks the steps: - A step that previously ended **COMPLETED** is **skipped** (not re-run). - A step that ended **FAILED**/**STOPPED** (or never ran) is **executed**, resuming from its persisted ExecutionContext where applicable. The two settings below tune this per step, configured on the `StepBuilder`. ## `startLimit(int)` - Bounds **how many times the step may be started** across ALL JobExecutions of the same JobInstance. - Default is effectively unlimited (`Integer.MAX_VALUE`). - Each time a restart actually *starts* that step, the count increments (it's tracked in step metadata). - When a start would exceed the limit, Spring Batch throws **`StartLimitExceededException`** and the job fails without running the step again. - **Use case**: a flaky step you don't want to retry forever — e.g. `startLimit(3)` means "give this step at most 3 attempts across restarts, then stop trying." ```java new StepBuilder("callFlakyService", repo) .<In, Out>chunk(100, tx) .reader(r).writer(w) .startLimit(3) // at most 3 starts across restarts .build(); ``` ## `allowStartIfComplete(boolean)` - Default **false** → a COMPLETED step is skipped on restart. - Set **true** → the step **always runs on restart**, even though it previously COMPLETED. - **Use cases**: steps that must run every time regardless of prior success — a pre-flight validation, a directory/staging setup, a cleanup/teardown step, or a step over a resource that isn't safely resumable and should just be redone. ```java new StepBuilder("validateInput", repo) .tasklet(validateTasklet, tx) .allowStartIfComplete(true) // re-run on every restart .build(); ``` ## How they interact - `allowStartIfComplete(true)` makes a completed step *eligible to run again*; `startLimit` still **counts each of those starts**. So a step that is both `allowStartIfComplete(true)` and `startLimit(3)` can be re-run on restarts but will throw `StartLimitExceededException` on the 4th start. - Both are strictly **within a single JobInstance**. Different JobInstances (different identifying parameters) reset all counters — it's a fresh instance. - Both operate **below** the job-level `restartable` flag: if the job is non-restartable, you never get to a second execution to exercise these at all. ## Gotchas - `startLimit` counts **starts**, not failures. Re-running an `allowStartIfComplete` step on successful restarts still consumes the budget. - Marking a chunk step `allowStartIfComplete(true)` means it **reprocesses from the beginning** each restart (its prior completion is ignored) — make sure that's idempotent. - These are step properties; don't confuse them with skip/retry limits (`faultTolerant().skipLimit()/retryLimit()`), which govern item-level fault tolerance *within* a single step execution, not restart across executions. - On restart, a step that FAILED resumes via ExecutionContext; `allowStartIfComplete` is irrelevant for it (it only matters for COMPLETED steps). ## When to use - `allowStartIfComplete(true)`: idempotent validation/setup/cleanup that must run each attempt. - `startLimit(n)`: bound retries on a step prone to repeated environmental failure so restarts don't loop forever.

  • How is startLimit different from a fault-tolerant step's retryLimit?
    startLimit bounds how many times a step is STARTED across restarts of a JobInstance (JobExecution-level, across separate launches). retryLimit (with faultTolerant()) governs retrying individual items/chunks WITHIN one step execution after exceptions. Different scopes entirely.
  • If a chunk step is marked allowStartIfComplete(true), does it resume from the last chunk on restart?
    No — allowStartIfComplete only applies to steps that already COMPLETED, and it makes them run again from the start, ignoring prior completion. Resume-from-last-chunk is for FAILED/STOPPED steps via their ExecutionContext. So an always-restart step should be idempotent.

saying these in an interview costs you the question

  • Saying startLimit counts failures rather than starts
  • Confusing startLimit/allowStartIfComplete (restart across executions) with skipLimit/retryLimit (fault tolerance within an execution)
  • Claiming allowStartIfComplete makes a step resume from its last chunk (it re-runs from the beginning)
  • Thinking these settings can override a job whose restartable flag is false

context