skip to content

How do getRunningExecutions and restart work in JobOperator, and how would you use them to safely manage a repeating job?

level: seniorimportance: should knowfreq 38%

answer

  1. getRunningExecutions → Set<Long> of in-progress ids
  2. STARTING/STARTED/STOPPING count as running
  3. Guard against overlap, but not a distributed lock
  4. restart = same instance (resume failure)
  5. restart refuses completed/running instances

basics

~20 s

getRunningExecutions(jobName) returns the ids of executions currently in flight for that job, so you can avoid launching overlapping runs. restart(executionId) re-runs a failed or stopped execution of the same instance, picking up per its restart semantics.

solid answer

~50 s

getRunningExecutions(String jobName) returns a Set<Long> of execution ids that are currently active — STARTING, STARTED, or STOPPING — for that job name (empty if none; NoSuchJobException for an unknown name). Operators use it as a guard: before starting a new run, check the set is empty to prevent concurrent overlapping executions of a job that isn't meant to run twice at once. restart(long executionId) takes a failed or stopped execution and launches a new execution of the *same JobInstance* (same identifying parameters). It refuses if the instance already completed (JobInstanceAlreadyCompleteException), if the execution is still running (JobExecutionAlreadyRunningException), or if the job/instance can't be found. A typical management flow: query getRunningExecutions to confirm nothing is in flight; if a prior run failed, restart its execution; otherwise startNextInstance for a fresh run. This gives an operator console safe start/resume behavior driven purely by names and ids.

code

java · 12 lines
java
public Long manage(String jobName, Long lastFailedExecId) throws Exception {
    Set<Long> running = jobOperator.getRunningExecutions(jobName);
    if (!running.isEmpty()) {
        throw new IllegalStateException(jobName + " already running: " + running);
    }
    if (lastFailedExecId != null) {
        // resume the SAME instance that failed/stopped
        return jobOperator.restart(lastFailedExecId);
    }
    // otherwise kick off a fresh instance (requires an incrementer)
    return jobOperator.startNextInstance(jobName);
}

go deeper

for a junior

Knows getRunningExecutions lists in-flight ids and restart re-runs a failed execution.

for a middle

Explains the running statuses and that restart targets the same instance/params.

for a senior

Builds a check-then-start/restart routine and knows the exceptions and the overlap-race caveat.

for a principal

Layers real distributed locking, distinguishes metadata guards from concurrency control, and ties orphaned-STARTED handling to abandon.

**getRunningExecutions.** Signature: `Set<Long> getRunningExecutions(String jobName)`. It consults the `JobRepository`/`JobExplorer` for executions of that job name whose status is *in progress* — `STARTING`, `STARTED`, or `STOPPING` (i.e., `BatchStatus.isRunning()`). It returns their execution ids. If the job name is unknown it throws `NoSuchJobException`; if the job exists but nothing is running it returns an empty set. Primary use: **overlap prevention**. Many batch jobs (nightly consolidation, a report build) must not run two instances simultaneously. Before launching, an operator/console checks `getRunningExecutions(jobName).isEmpty()` and only proceeds if so. Note this is a check on Spring Batch metadata, not a distributed lock — in a multi-node deployment two nodes could both read 'empty' and both start; if you need hard mutual exclusion across nodes you layer a real lock (DB row lock, ShedLock, etc.) on top. The metadata check is still valuable within a single instance and as a first-line guard. **restart.** Signature: `Long restart(long executionId)`. Given the id of a previous execution that ended `FAILED` or `STOPPED`, it creates a *new* `JobExecution` for the *same* `JobInstance` (same identifying parameters) and runs it. What happens inside the re-run — which steps re-execute, whether reader/writer state resumes — is governed by restart/resume semantics (step allow-start-if-complete, save-state, execution context) which live under a sibling topic; from the operator's standpoint the point is: `restart` continues the *same instance*, whereas `start`/`startNextInstance` create *new* instances. Failure modes: - `JobInstanceAlreadyCompleteException` — the instance already completed successfully; there's nothing to restart. - `JobExecutionAlreadyRunningException` — that execution is still running; you can't restart a live one. - `NoSuchJobExecutionException` / `NoSuchJobException` — bad id or job not registered. - `JobRestartException` — the instance is not in a restartable state. **Putting it together — a safe management routine.** 1. `Set<Long> running = jobOperator.getRunningExecutions("nightlyJob");` 2. If `running` is non-empty → refuse to start (a run is in flight); optionally surface those ids in the console. 3. Else decide: was the last run a failure? If you hold that execution id, call `restart(id)` to resume the same instance. If the last run succeeded and you want a new run, call `startNextInstance("nightlyJob")` (needs an incrementer). **Gotchas.** - getRunningExecutions is a snapshot; there's a race between checking and starting — hence not a substitute for a real lock across nodes. - restart is not the same as startNextInstance: restart = same instance/params (resume a failure), startNextInstance = brand-new instance (next scheduled run). - An orphaned STARTED execution (crashed JVM) will appear 'running' forever in metadata and block restart; that's the abandon use case.

  • Is getRunningExecutions a reliable way to guarantee only one run cluster-wide?
    No. It's a metadata snapshot with a check-then-act race; across multiple nodes both could see 'empty' and start. Add a real distributed lock for hard mutual exclusion.
  • How is restart different from startNextInstance?
    restart resumes the SAME JobInstance (same identifying parameters) of a failed/stopped run; startNextInstance creates a NEW instance via the incrementer. One continues, the other begins afresh.
  • Why might restart throw JobExecutionAlreadyRunningException?
    Because the target execution (or its instance) is still marked running — often an orphaned STARTED execution from a crash that must be abandoned first.

saying these in an interview costs you the question

  • Treating getRunningExecutions as a cluster-wide lock
  • Confusing restart (same instance) with startNextInstance (new instance)
  • Thinking a completed instance can be restarted

context