skip to content

What design constraints (statelessness, idempotency, side effects) apply to an ItemProcessor in a fault-tolerant or multi-threaded step, and why?

level: principalimportance: should knowfreq 30%

answer

  1. Retry re-processes items -> not exactly-once
  2. Idempotent, no external side effects
  3. Side effects belong in the writer
  4. Multi-thread shares one bean -> no mutable fields
  5. Stateful processor needs ItemStream + single thread

basics

~20 s

Keep processors stateless, side-effect-free, and idempotent. In fault-tolerant steps a chunk rollback causes items to be re-processed, so process() may be called more than once per item; in multi-threaded steps the same instance runs concurrently, so mutable fields are unsafe.

solid answer

~40 s

Two step features drive these constraints. First, fault tolerance: when a chunk fails and Spring Batch retries or scans-to-skip, it rolls back the transaction and re-reads and re-processes items — so process() is NOT guaranteed exactly-once. It must be idempotent and free of committed side effects (don't send emails, call non-idempotent APIs, or mutate shared state inside it), otherwise a retry double-applies them. Second, multi-threaded/partitioned steps share one processor bean across threads, so any mutable instance state is a race condition; keep processors immutable/stateless or use ThreadLocal. Prefer pure input->output transformation. If a processor genuinely needs to accumulate state (rare), it must implement ItemStream and be non-thread-safe-by-design, which pushes you toward a single-threaded step. Also account for the retry re-entry when doing lookups or caching. These rules keep jobs correct and restartable.

code

java · 21 lines
java
// GOOD: stateless, idempotent, thread-safe (only reads from a thread-safe cache)
public class PriceProcessor implements ItemProcessor<Order, PricedOrder> {
    private final RateLookup rates; // thread-safe, read-only during the run
    public PriceProcessor(RateLookup rates) { this.rates = rates; }
    @Override public PricedOrder process(Order o) {
        BigDecimal total = o.getAmount().multiply(rates.forCurrency(o.getCurrency()));
        return new PricedOrder(o.getId(), total); // pure transform
    }
}

// BAD: mutable running state (race in multi-threaded step) + external side effect (duplicated on retry)
public class BadProcessor implements ItemProcessor<Order, Order> {
    private long seen = 0;                 // shared mutable state -> data race
    private final MailClient mail;
    public BadProcessor(MailClient mail) { this.mail = mail; }
    @Override public Order process(Order o) {
        seen++;                            // corrupt under concurrency
        mail.send(o.getEmail());           // re-sent on chunk retry (not idempotent)
        return o;
    }
}

go deeper

for a junior

Know processors should be simple and not do database writes.

for a middle

Understand statelessness and that side effects belong in the writer.

for a senior

Explain retry re-processing (no exactly-once) and thread-safety of the shared bean.

for a principal

Reason about idempotency, restartability via ExecutionContext, and when statefulness forces single-threaded design or an outbox.

**Why the constraints exist — two mechanisms.** 1. **Fault tolerance and retry/skip re-processing.** In a `.faultTolerant()` chunk step, when the write (or process) of a chunk throws a retryable/skippable exception, Spring Batch **rolls back** the chunk transaction. To honor skip/retry semantics it then **re-reads and re-processes** the items (often one at a time on the retry/scan pass). Consequences: - `ItemProcessor.process()` can be invoked **more than once for the same logical item**. There is no exactly-once guarantee at the processor. - Therefore the processor must be **idempotent**: running it twice on the same input must yield the same output and cause no cumulative effect. - It must avoid **externally-visible side effects** — sending a message, calling a payment API, incrementing a counter in another system — because those would be duplicated on retry. Side effects belong in the `ItemWriter` (which participates in the chunk transaction) or must themselves be idempotent. - There is a config nuance: by default the processor result is *not* cached across retries; `SimpleChunkProcessor` re-runs the processor unless `processorTransactional`/reprocessing settings say otherwise — so assume re-execution. 2. **Concurrency — multi-threaded and partitioned steps.** A step with a `TaskExecutor` (`.taskExecutor(...)`) or partitioning shares a **single processor bean instance** across worker threads (unless step-scoped per-partition). Any mutable **instance field** (a running total, a cached last-seen key, a non-thread-safe collection) becomes a data race producing corrupt or non-deterministic results. Hence processors should be **stateless/immutable**; if per-invocation scratch state is unavoidable, keep it local to the method or use `ThreadLocal`. **Design guidance.** - Model the processor as a **pure function** `O f(I)` — no reads/writes of shared mutable state, no I/O with side effects. Injected read-only collaborators (services, caches, repositories used for lookups) are fine as long as they are thread-safe and the lookups are safe to repeat. - Do **enrichment lookups** carefully: caching is fine if the cache is thread-safe (e.g. `ConcurrentHashMap`/`Caffeine`) and the underlying data is stable during the run. - Put **transactional side effects** (persistence, message publication via a transactional resource) in the writer, so rollback undoes them; do not perform them in the processor. - If the processor *must* be stateful and ordered (e.g. running aggregation), you generally cannot make the step multi-threaded, and you should implement `ItemStream` to persist/restore that state in the `ExecutionContext` for **restartability** — but this is a design smell worth questioning. **Restartability angle.** Batch jobs are restartable from the last committed chunk. Hidden state in a processor that isn't saved to the `ExecutionContext` will be lost on restart, causing incorrect results. Statelessness sidesteps the whole problem. **Gotchas summary.** - Assuming exactly-once processing (wrong under retry). - Sending emails / calling external non-idempotent APIs from `process()`. - Mutable fields in a processor used by a multi-threaded step. - Relying on processing order in a multi-threaded step (order is not guaranteed). **When these matter most.** High-volume jobs that enable fault tolerance and/or parallelism — exactly the jobs where correctness bugs are hardest to reproduce. For a simple single-threaded, non-fault-tolerant step the constraints relax, but coding to them anyway keeps the step safe to scale later.

  • Is ItemProcessor.process() guaranteed to run exactly once per item?
    No. In a fault-tolerant step a chunk rollback re-reads and re-processes items, so process() can run multiple times for the same item. Design it to be idempotent.
  • You must send a notification per processed record. Where should that go?
    Not in the processor — it would duplicate on retry. Put it in the writer (transactional so rollback undoes it) or make it idempotent, or emit it via an outbox written in the same chunk transaction.
  • Why are mutable fields dangerous in a processor used by a multi-threaded step?
    Spring Batch shares one processor instance across worker threads, so mutable instance state is a data race yielding corrupt/non-deterministic results and breaks restartability.

saying these in an interview costs you the question

  • Assuming exactly-once processing semantics
  • Performing external side effects (emails, non-idempotent API calls) inside process()
  • Keeping running totals in processor instance fields for a multi-threaded step
  • Assuming processing order is preserved under concurrency
  • Believing hidden processor state survives a job restart

context