Design the launch/idempotency strategy for a nightly financial reconciliation job that must be safe to re-trigger by an operator, resumable after a crash, and must never double-post entries. What are the trade-offs?
answer
- businessDate = identity + correctness boundary
- completed->reject, failed->restart, running->reject
- no RunIdIncrementer (would allow double-post)
- restartable readers save ExecutionContext
- idempotent postings: natural-key upsert / dedupe marker
basics
~20 sKey the job by business date (identifying param). Same date + completed = rejected (no double-run); same date + failed = restart and resume. Make writers idempotent via upserts so replayed chunks never double-post. Reserve an incrementer only for deliberately advancing to a new date.
solid answer
~50 sI'd anchor the JobInstance identity on an identifying `businessDate` parameter. That gives three behaviors for free from the JobRepository: relaunching a COMPLETED date throws JobInstanceAlreadyCompleteException (operator double-click is safely rejected); relaunching a FAILED/STOPPED date restarts the same instance and resumes from the last committed chunk; a concurrent launch throws JobExecutionAlreadyRunningException. For crash resumability I ensure restartable readers persist their position in the ExecutionContext and that chunk business writes share the batch transaction where possible. To guarantee no double-posting, ledger writes are idempotent — upsert on a natural entry key or an atomic 'already-posted' dedupe marker — because restart and retry give at-least-once semantics. I'd avoid RunIdIncrementer here since always-advancing run ids would let two runs post the same date; the business-date key is the correctness boundary. Advancing to the next date is an explicit, separate action.
code
java · 38 lines@Bean
Job reconciliationJob(JobRepository repo, Step postEntriesStep) {
// NO incrementer: identity is the businessDate parameter, which is the
// correctness boundary against double-posting a given day.
return new JobBuilder("reconciliationJob", repo)
.start(postEntriesStep)
.build();
}
@Bean
Step postEntriesStep(JobRepository repo, PlatformTransactionManager tx,
ItemReader<Txn> reader, ItemWriter<LedgerEntry> writer) {
return new StepBuilder("postEntriesStep", repo)
.<Txn, LedgerEntry>chunk(500, tx) // chunk = restart granularity
.reader(reader) // restartable: saves position in ExecutionContext
.writer(writer) // idempotent upsert (see below)
.faultTolerant()
.build();
}
@Bean
JdbcBatchItemWriter<LedgerEntry> writer(DataSource ds) {
return new JdbcBatchItemWriterBuilder<LedgerEntry>()
.dataSource(ds)
.sql("""
INSERT INTO ledger (entry_key, business_date, account_id, amount)
VALUES (:entryKey, :businessDate, :accountId, :amount)
ON CONFLICT (entry_key) DO NOTHING -- replay is a no-op, never double-posts
""")
.beanMapped()
.build();
}
// Launch: same businessDate re-trigger is guarded by the JobRepository.
JobParameters p = new JobParametersBuilder()
.addString("businessDate", "2026-07-21") // identifying
.toJobParameters();
// completed -> JobInstanceAlreadyCompleteException; failed -> resumes; running -> already-running.go deeper
Can state 'don't run twice' but not architect the guarantees.
Picks business-date keying and upserts but may miss the cross-system and concurrency edges.
Combines identity guard, restartable readers, and idempotent writers with clear reasoning.
Audits every failure mode, justifies rejecting the incrementer, addresses external sinks, and defines the operational runbook and monitoring.
**Requirements decomposed.** (1) Operator can safely re-trigger — no accidental duplicate posting. (2) Resumable after a crash. (3) Never double-post ledger entries. These map to three Spring Batch mechanisms: JobInstance identity, restart semantics, and writer idempotency. **1. Identity: business-date keying.** Make `businessDate` (the reconciliation window, e.g. `2026-07-21`) an **identifying** job parameter. Consequences from the JobRepository guard: - Re-launch same date after **COMPLETED** → `JobInstanceAlreadyCompleteException`. This *is* the anti-double-run control: an operator hitting 'run' twice, or a scheduler firing twice, is rejected by the framework. - Re-launch same date after **FAILED/STOPPED** → restart: a new JobExecution on the *same* JobInstance, resuming restartable steps. - Re-launch while **running** → `JobExecutionAlreadyRunningException` (guard against overlap). This is why I would **not** use `RunIdIncrementer` for this job: an always-incrementing run.id would make every trigger a new instance, so two runs could both post `2026-07-21` — exactly the failure we must prevent. The business date is the correctness boundary. **2. Resumability.** For a crash mid-run to resume rather than restart-from-zero: - Use **restartable ItemReaders** that persist their cursor/offset into the step `ExecutionContext` (e.g. `JdbcCursorItemReader`/`JdbcPagingItemReader`, flat-file readers that save line count). Spring Batch reloads that context on restart. - Keep chunk business writes in the **same transaction** as the batch metadata where the sink is the same datasource, so a chunk commit is atomic with the bookkeeping — minimizing the replay window. - Set a sensible **commit-interval (chunk size)**: smaller chunks = less rework on restart but more transaction overhead. - Consider `allowStartIfComplete(false)` (default) so completed steps are skipped on restart, and be deliberate about `startLimit`. **3. No double-posting: idempotent writes.** Restart and chunk retry/skip give **at-least-once** semantics — a chunk may be applied twice. So ledger posting must be idempotent: - **Natural key + upsert.** Each entry has a deterministic key (e.g. `hash(businessDate, accountId, txnId)`); write via `INSERT ... ON CONFLICT DO NOTHING/UPDATE` or `MERGE`. Replays re-apply the same entry, never a second row. - **Atomic dedupe marker.** If posting is to an external ledger/API, record a 'posted' marker keyed by the entry key in the same local transaction that triggers the effect, or use a downstream idempotency key. Check-then-post is unsafe under concurrency unless the check-and-mark is atomic. - **Avoid DB-generated ids** for postings; they make inserts non-idempotent. **Operational trade-offs.** - *Business-date key vs incrementer:* the date key preserves the double-run guard but means 'run it again for real' requires clearing/failing the prior instance or moving to the next date — a deliberate operator action. An incrementer trades that safety for effortless repetition. - *Restart vs re-run from scratch:* resuming is faster and avoids re-posting, but requires readers/writers designed for it; a non-restartable reader forces from-zero replay, which is only safe because writers are idempotent. - *Chunk size:* correctness is unaffected, but it tunes restart cost vs throughput. - *Same-datasource transaction:* strongest atomicity, but cross-system reconciliation often can't share a transaction, pushing you toward idempotency keys and eventual-consistency reasoning. **Failure-mode audit.** Walk each: crash before any commit (restart re-reads, no effect); crash after chunk commit but before metadata (idempotent upsert absorbs the replay); operator double-click on completed date (rejected by guard); scheduler double-fire while running (rejected as already-running); partial external post (dedupe marker/idempotency key prevents re-charge). If all five are safe, the design holds. **Governance.** Add monitoring on `BATCH_JOB_EXECUTION` statuses, alert on repeated FAILED for a date, and give operators a controlled 'restart date X' vs 'abandon and advance' runbook rather than ad-hoc parameter fiddling.
- Why not just use RunIdIncrementer so operators can always re-run without an exception?Because an ever-advancing run.id makes each trigger a new JobInstance, so two runs could both post the same business date — the exact double-post we must prevent. The business-date key intentionally trades effortless re-runs for the framework-level guarantee that a completed date can't run twice.
- The reconciliation reads from one datasource and posts to a separate ledger system that can't join the batch transaction. How do you still avoid double-posting?Use an idempotency key the ledger honors (so it dedupes replays), or keep a local 'posted' table keyed by entry key that you write atomically before/with the post and check on replay. Without a shared transaction you rely on downstream idempotency plus your dedupe store, and design for at-least-once.
- An operator genuinely needs to re-run a date that already COMPLETED (bad source data was fixed). How do you handle it safely?Provide a controlled runbook: either roll back/void the prior postings and start a fresh JobInstance under a new run key, or introduce a corrective re-run parameter. Because writers are idempotent upserts keyed by entry, a deliberate re-post overwrites cleanly rather than duplicating — but this must be an explicit, audited action, not an accidental relaunch.
saying these in an interview costs you the question
- Slapping RunIdIncrementer on a job that must not double-post a given period
- Assuming restart alone prevents duplicate side effects without idempotent writers
- Treating cross-system posts as if they share the batch transaction
- Letting operators re-run completed periods via ad-hoc parameter changes with no dedupe