skip to content

You must trigger a long-running batch job from a REST endpoint. Design the launching mechanism and justify the trade-offs.

level: principalimportance: should knowfreq 35%

answer

  1. never run() on the HTTP thread (sync blocks)
  2. bounded ThreadPoolTaskExecutor = concurrency policy
  3. 202 + execution id, poll status via JobExplorer
  4. incrementer for independence vs stable params for 409 dedup
  5. orphaned STARTED executions on crash -> recovery

basics

~20 s

Don't use the default blocking launcher on an HTTP thread. Configure a TaskExecutorJobLauncher with a bounded ThreadPoolTaskExecutor so run() returns immediately with an execution id; the client polls a status endpoint that reads the JobRepository.

solid answer

~50 s

The core problem: the default TaskExecutorJobLauncher uses a SyncTaskExecutor, so run() blocks the HTTP worker for the whole job — unacceptable for a long job. I'd configure the launcher with a bounded ThreadPoolTaskExecutor: run() then persists the JobExecution, hands off to a pool thread, and returns immediately with the execution id in STARTING/STARTED. The endpoint returns 202 Accepted + the id; a separate GET reads status from JobExplorer/JobRepository. I'd bound concurrency (pool + queue sizes) so a burst of requests can't exhaust threads, and add a JobParametersIncrementer or unique parameter so repeated calls create distinct instances (avoiding JobInstanceAlreadyCompleteException); if I want dedup I'd instead reuse identifying params and let JobExecutionAlreadyRunningException reject concurrent duplicates as 409. I'd also plan graceful shutdown: in-flight jobs and the metadata store must be consistent, and orphaned STARTED executions recovered on restart.

code

java · 27 lines
java
@RestController
@RequestMapping("/jobs/import")
class ImportJobController {
    private final JobLauncher asyncJobLauncher; // backed by ThreadPoolTaskExecutor
    private final Job importJob;
    private final JobExplorer jobExplorer;

    ImportJobController(JobLauncher asyncJobLauncher, Job importJob, JobExplorer jobExplorer) {
        this.asyncJobLauncher = asyncJobLauncher; this.importJob = importJob; this.jobExplorer = jobExplorer;
    }

    @PostMapping
    ResponseEntity<?> start() throws Exception {
        var params = new JobParametersBuilder()
                .addLong("run.id", System.currentTimeMillis()) // new instance per call
                .toJobParameters();
        var exec = asyncJobLauncher.run(importJob, params); // returns STARTING/STARTED
        return ResponseEntity.accepted().body(exec.getId());   // 202 + id to poll
    }

    @GetMapping("/{id}")
    ResponseEntity<String> status(@PathVariable long id) {
        var exec = jobExplorer.getJobExecution(id);
        return exec == null ? ResponseEntity.notFound().build()
                            : ResponseEntity.ok(exec.getStatus().name());
    }
}

go deeper

for a junior

Recognize you shouldn't block the request thread for a long job.

for a middle

Propose an async launcher and a status-polling endpoint.

for a senior

Bound the pool, choose an instance-identity policy, map batch exceptions to HTTP codes, and use JobExplorer for status.

for a principal

Reason end-to-end: backpressure/rejection policy, graceful shutdown + orphan recovery, dedup vs independence, and when to move to a separate worker/container for scaling.

## Why the naive approach is wrong Injecting the default `JobLauncher` and calling `run()` inside a `@PostMapping` **ties up the servlet/worker thread for the entire job**, because the default `TaskExecutorJobLauncher` runs on a `SyncTaskExecutor` (caller thread). For a job measured in minutes/hours this blocks the connection, exhausts the web thread pool under load, and gives the client no id to track. Async launching is mandatory. ## The design ### 1. Async, bounded launcher ```java @Bean JobLauncher asyncJobLauncher(JobRepository repo) throws Exception { var pool = new ThreadPoolTaskExecutor(); pool.setCorePoolSize(2); pool.setMaxPoolSize(4); pool.setQueueCapacity(20); pool.setThreadNamePrefix("batch-launch-"); pool.initialize(); var launcher = new TaskExecutorJobLauncher(); launcher.setJobRepository(repo); launcher.setTaskExecutor(pool); launcher.afterPropertiesSet(); return launcher; } ``` Choose `ThreadPoolTaskExecutor` (bounded) over `SimpleAsyncTaskExecutor` (unbounded, thread-per-task) so a spike of requests can't create unbounded threads. The pool + queue **is** your concurrency policy: it caps how many jobs run at once and how many wait. ### 2. Endpoint contract - `POST /jobs/import` builds `JobParameters`, calls `run()`, and returns **202 Accepted** with the `JobExecution` id (and a `Location` to the status resource). - `GET /jobs/import/{executionId}` reads status via `JobExplorer` / `JobOperator` and returns COMPLETED/FAILED/STARTED. This is the poll-for-result pattern async launching forces, because the returned status right after `run()` is non-terminal (STARTING/STARTED). ### 3. Instance-identity policy (idempotency vs. dedup) Two opposite choices, both driven by `JobParameters` identity: - **Every request is independent:** add a `RunIdIncrementer` or a unique parameter (uuid/timestamp) so each call is a new `JobInstance`; you never hit `JobInstanceAlreadyCompleteException`. - **Dedup concurrent duplicates:** keep identifying params stable (e.g. a business key) so a second concurrent call to the same instance throws `JobExecutionAlreadyRunningException`, which you map to **409 Conflict** — natural request de-duplication. A completed one throws `JobInstanceAlreadyCompleteException` (map to 409/'already done'). ### 4. Backpressure & overload With a bounded pool, when the queue is full the executor's rejection policy kicks in (default `AbortPolicy` -> `TaskRejectedException`). Decide: reject the request with **429/503**, or use `CallerRunsPolicy` (dangerous here — it would run on the HTTP thread, reintroducing blocking). Explicitly configure this; don't leave it to chance. ### 5. Lifecycle & recovery - **Graceful shutdown:** set the pool to wait for tasks and a timeout; but a long job may outlive shutdown. On abrupt termination, executions left in `STARTED` become **orphans**. Plan a startup recovery step (e.g. `JobRepository`/`JobExplorer` scan) to mark abandoned executions or use `JobOperator` to handle them. - **Transactions:** the `JobRepository` persists the JobExecution *before* the async thread runs, so the metadata is durable even though the work is async — good for tracking and restart. - **Observability:** expose metrics (Micrometer batch metrics), and thread-name prefixes for the pool so async jobs are traceable. ### 6. When to push further If jobs are heavy or the app is horizontally scaled, in-JVM async launching doesn't distribute load and a redeploy kills in-flight jobs. At that scale prefer a **separate batch worker** (a job app started per run, e.g. a k8s Job / cron using `JobLauncherApplicationRunner` + exit codes) or remote partitioning, with the web tier only enqueuing requests. The REST-triggered async launcher is the right middle ground for moderate, single-instance workloads. ## Trade-off summary - Sync launcher: simple, correct status immediately, but blocks — only for CLI/startup. - Async in-JVM (ThreadPoolTaskExecutor): non-blocking endpoint, bounded concurrency, needs poll-for-status + orphan recovery + rejection policy. - Separate worker/container: best isolation and scaling, more moving parts and orchestration.

  • How would you dedup two identical concurrent requests instead of running both?
    Keep the identifying JobParameters stable (a business key). The second launch of the same running instance throws JobExecutionAlreadyRunningException, which you map to HTTP 409 — natural de-duplication. Vary a param only when you want independent runs.
  • The app is killed mid-job. What state is left and how do you recover?
    The JobExecution stays STARTED (orphaned) in the JobRepository. On startup, scan for abandoned executions and mark/abandon them (via JobOperator/JobRepository) before relaunching, so restart logic isn't confused by phantom running executions.
  • Why prefer ThreadPoolTaskExecutor over SimpleAsyncTaskExecutor here?
    SimpleAsyncTaskExecutor is thread-per-task and unbounded by default, so a request spike can exhaust threads. A bounded ThreadPoolTaskExecutor caps concurrency and queues/reject excess, giving predictable resource use and backpressure.

saying these in an interview costs you the question

  • Calling the default (sync) launcher directly in the controller and blocking the HTTP thread.
  • Using an unbounded SimpleAsyncTaskExecutor with no concurrency cap.
  • Returning the job's final status synchronously after an async run() (it's still STARTING/STARTED).
  • Using CallerRunsPolicy on the launcher pool (pushes work back onto the HTTP thread).
  • Ignoring orphaned STARTED executions after a crash.

context