When an exception is registered for both retry and skip, how do retry and skip interact, and how do their counters and the writer scan interplay?
answer
- Order: retry -> skip -> fail
- skip counts once, only after retries exhausted
- writer scan isolates the item to skip
- skipLimit is cumulative across the step
- retryable-but-not-skippable => step fails when exhausted
basics
~20 sRetry runs first: Batch retries the item up to retryLimit. Only when retries are exhausted and the same exception is skippable does it skip the item, counting it toward skipLimit. If skipLimit is exceeded the step fails. So the order is retry, then skip, then fail.
solid answer
~50 sRetry and skip are layered, not alternatives. For an exception registered with both retry() and skip(), Batch first exhausts the retry policy (retryLimit / RetryPolicy) for that item. If it still fails and the exception is skippable, the item is skipped and the skip counter increments; when total skips exceed skipLimit the step fails. During the writer phase this plays out through scan mode: after retries are exhausted on a batched write, Batch re-drives the chunk item-by-item to isolate the failing item, so only the genuinely bad item is skipped while its chunk-mates commit. Two subtleties: each retry attempt does NOT increment the skip count — only the final give-up does; and an exception that is retryable but not skippable will fail the step once retries run out. Classify deliberately: retry transient errors, skip poison records, and reserve non-skippable/non-retryable for truly fatal conditions.
code
java · 17 linesreturn new StepBuilder("step", jobRepository)
.<In, Out>chunk(50, tx)
.reader(reader).processor(processor).writer(writer)
.faultTolerant()
// transient: retry a few times, and if it still fails for one record, skip it
.retry(OptimisticLockingFailureException.class).retryLimit(3)
.skip(OptimisticLockingFailureException.class)
// poison data: never retry, skip immediately
.skip(FlatFileParseException.class)
.noRetry(FlatFileParseException.class)
.skipLimit(20) // cumulative across the whole step
.listener(new SkipListener<In, Out>() {
@Override public void onSkipInWrite(Out item, Throwable t) {
log.warn("skipped after retries exhausted: {}", item, t);
}
})
.build();go deeper
Know that retry happens first and skip is the fallback when retries run out.
Explain that skip counts once after retries exhaust and skipLimit failing the step.
Tie in writer scan-mode isolation and deliberate exception classification (retry vs skip vs both).
Design the retry/skip/fail escalation with cumulative-budget awareness, monitoring, and idempotency, distinguishing systemic downstream failure from poison data.
### Retry and skip are layers, applied in order Think of three escalating outcomes for a failing item: **retry -> skip -> fail**. When an exception type is registered with *both* `retry(X.class)` and `skip(X.class)`: 1. **Retry first.** Batch applies the retry policy (attempts up to `retryLimit`). Each attempt re-drives the item within a rolled-back chunk. 2. **Skip on exhaustion.** If, after the last allowed attempt, the item still throws and the exception is **skippable**, Batch **skips** the item — it is excluded from output and the **skip count** increments. Skip listeners (`SkipListener.onSkipInWrite/InProcess/InRead`) fire here. 3. **Fail on skip overflow.** When the cumulative skip count exceeds **`skipLimit`**, the step fails with `SkipLimitExceededException`. ### Counters do not double-count A retried attempt is **not** a skip. The skip counter increments **once**, at the moment Batch gives up retrying and skips. So `retryLimit(3)` + `skipLimit(10)` means each poison item burns up to 3 attempts and then costs exactly **1** against the skip budget. ### How the writer scan ties it together Because a batched write fails as a unit, after retries are exhausted Batch enters **scan mode** and re-processes/re-writes the chunk **one item at a time**. This is essential for retry-then-skip to be precise: without single-item isolation, Batch could not know *which* item to skip, and would have to fail or skip the whole chunk. Scan mode lets the good items commit while only the offending item is skipped. This also means processors/writers may run several times per item — idempotency again matters. ### Classification is a design decision - **Retryable only** (not skippable): transient error that must eventually succeed or fail the job (e.g. a mandatory downstream). Retries exhausted → **step fails**. - **Skippable only** (not retryable): deterministic poison data (parse/validation). No point retrying → skipped immediately. - **Both:** a transient error that, if it stubbornly persists for one record, should not sink the whole job → retry a few times, then skip and move on. - Use `noRetry(...)`/`noSkip(...)` to carve exceptions out of broader registrations, and `noRollback(...)` for exceptions that should not roll the chunk back at all. ### Gotchas at principal level - **Ordering of exhaustion:** retries must be exhausted **before** a skip is considered — candidates often invert this. - **skipLimit is cumulative across the whole step**, not per chunk; a wave of transient failures can silently eat the budget and then fail late. - **SkipListener runs in its own transaction** after the skip, so a listener failure has separate semantics from the item. - **Metrics/alerting:** high skip counts driven by retry exhaustion usually signal a systemic downstream problem, not bad data — monitor retries and skips separately. - Retry and skip policies can be replaced wholesale with custom `RetryPolicy`/`SkipPolicy` when the count-based defaults are too blunt.
- If an exception is retryable but not skippable, what happens when retries are exhausted?The step fails — there is no skip fallback, so the exception propagates after the last retry. The item cannot be dropped; retry is the only mitigation and its exhaustion is terminal for that step run.
- Does each retry attempt consume the skip budget?No. Retries are counted by the retry policy per item; the skip counter increments exactly once, only at the moment Batch stops retrying and skips the item. retryLimit and skipLimit are independent budgets.
- Why is writer scan mode necessary for retry-then-skip to work correctly?A batched write fails as a unit, so Batch cannot attribute the failure to a specific item. Scan mode re-drives the chunk one item at a time, isolating the culprit so it can be skipped while the rest of the chunk commits.
saying these in an interview costs you the question
- Saying skip happens before retries are exhausted
- Claiming each retry attempt increments the skip count
- Thinking a retryable-but-not-skippable error is silently dropped
- Believing skipLimit is per-chunk rather than cumulative across the step
- Assuming the whole chunk is skipped rather than the single isolated item