skip to content

Design the launch/idempotency strategy for a nightly financial reconciliation job that must be safe to re-trigger by an operator, resumable after a crash, and must never double-post entries. What are the trade-offs?

level: principalimportance: should knowfreq 35%

answer

  1. businessDate = identity + correctness boundary
  2. completed->reject, failed->restart, running->reject
  3. no RunIdIncrementer (would allow double-post)
  4. restartable readers save ExecutionContext
  5. idempotent postings: natural-key upsert / dedupe marker

basics

~20 s

Key the job by business date (identifying param). Same date + completed = rejected (no double-run); same date + failed = restart and resume. Make writers idempotent via upserts so replayed chunks never double-post. Reserve an incrementer only for deliberately advancing to a new date.

solid answer

~50 s

I'd anchor the JobInstance identity on an identifying `businessDate` parameter. That gives three behaviors for free from the JobRepository: relaunching a COMPLETED date throws JobInstanceAlreadyCompleteException (operator double-click is safely rejected); relaunching a FAILED/STOPPED date restarts the same instance and resumes from the last committed chunk; a concurrent launch throws JobExecutionAlreadyRunningException. For crash resumability I ensure restartable readers persist their position in the ExecutionContext and that chunk business writes share the batch transaction where possible. To guarantee no double-posting, ledger writes are idempotent — upsert on a natural entry key or an atomic 'already-posted' dedupe marker — because restart and retry give at-least-once semantics. I'd avoid RunIdIncrementer here since always-advancing run ids would let two runs post the same date; the business-date key is the correctness boundary. Advancing to the next date is an explicit, separate action.

code

java · 38 lines
java
@Bean
Job reconciliationJob(JobRepository repo, Step postEntriesStep) {
    // NO incrementer: identity is the businessDate parameter, which is the
    // correctness boundary against double-posting a given day.
    return new JobBuilder("reconciliationJob", repo)
            .start(postEntriesStep)
            .build();
}

@Bean
Step postEntriesStep(JobRepository repo, PlatformTransactionManager tx,
                     ItemReader<Txn> reader, ItemWriter<LedgerEntry> writer) {
    return new StepBuilder("postEntriesStep", repo)
            .<Txn, LedgerEntry>chunk(500, tx)   // chunk = restart granularity
            .reader(reader)                     // restartable: saves position in ExecutionContext
            .writer(writer)                     // idempotent upsert (see below)
            .faultTolerant()
            .build();
}

@Bean
JdbcBatchItemWriter<LedgerEntry> writer(DataSource ds) {
    return new JdbcBatchItemWriterBuilder<LedgerEntry>()
            .dataSource(ds)
            .sql("""
                INSERT INTO ledger (entry_key, business_date, account_id, amount)
                VALUES (:entryKey, :businessDate, :accountId, :amount)
                ON CONFLICT (entry_key) DO NOTHING  -- replay is a no-op, never double-posts
                """)
            .beanMapped()
            .build();
}

// Launch: same businessDate re-trigger is guarded by the JobRepository.
JobParameters p = new JobParametersBuilder()
        .addString("businessDate", "2026-07-21") // identifying
        .toJobParameters();
// completed -> JobInstanceAlreadyCompleteException; failed -> resumes; running -> already-running.

go deeper

for a junior

Can state 'don't run twice' but not architect the guarantees.

for a middle

Picks business-date keying and upserts but may miss the cross-system and concurrency edges.

for a senior

Combines identity guard, restartable readers, and idempotent writers with clear reasoning.

for a principal

Audits every failure mode, justifies rejecting the incrementer, addresses external sinks, and defines the operational runbook and monitoring.

**Requirements decomposed.** (1) Operator can safely re-trigger — no accidental duplicate posting. (2) Resumable after a crash. (3) Never double-post ledger entries. These map to three Spring Batch mechanisms: JobInstance identity, restart semantics, and writer idempotency. **1. Identity: business-date keying.** Make `businessDate` (the reconciliation window, e.g. `2026-07-21`) an **identifying** job parameter. Consequences from the JobRepository guard: - Re-launch same date after **COMPLETED** → `JobInstanceAlreadyCompleteException`. This *is* the anti-double-run control: an operator hitting 'run' twice, or a scheduler firing twice, is rejected by the framework. - Re-launch same date after **FAILED/STOPPED** → restart: a new JobExecution on the *same* JobInstance, resuming restartable steps. - Re-launch while **running** → `JobExecutionAlreadyRunningException` (guard against overlap). This is why I would **not** use `RunIdIncrementer` for this job: an always-incrementing run.id would make every trigger a new instance, so two runs could both post `2026-07-21` — exactly the failure we must prevent. The business date is the correctness boundary. **2. Resumability.** For a crash mid-run to resume rather than restart-from-zero: - Use **restartable ItemReaders** that persist their cursor/offset into the step `ExecutionContext` (e.g. `JdbcCursorItemReader`/`JdbcPagingItemReader`, flat-file readers that save line count). Spring Batch reloads that context on restart. - Keep chunk business writes in the **same transaction** as the batch metadata where the sink is the same datasource, so a chunk commit is atomic with the bookkeeping — minimizing the replay window. - Set a sensible **commit-interval (chunk size)**: smaller chunks = less rework on restart but more transaction overhead. - Consider `allowStartIfComplete(false)` (default) so completed steps are skipped on restart, and be deliberate about `startLimit`. **3. No double-posting: idempotent writes.** Restart and chunk retry/skip give **at-least-once** semantics — a chunk may be applied twice. So ledger posting must be idempotent: - **Natural key + upsert.** Each entry has a deterministic key (e.g. `hash(businessDate, accountId, txnId)`); write via `INSERT ... ON CONFLICT DO NOTHING/UPDATE` or `MERGE`. Replays re-apply the same entry, never a second row. - **Atomic dedupe marker.** If posting is to an external ledger/API, record a 'posted' marker keyed by the entry key in the same local transaction that triggers the effect, or use a downstream idempotency key. Check-then-post is unsafe under concurrency unless the check-and-mark is atomic. - **Avoid DB-generated ids** for postings; they make inserts non-idempotent. **Operational trade-offs.** - *Business-date key vs incrementer:* the date key preserves the double-run guard but means 'run it again for real' requires clearing/failing the prior instance or moving to the next date — a deliberate operator action. An incrementer trades that safety for effortless repetition. - *Restart vs re-run from scratch:* resuming is faster and avoids re-posting, but requires readers/writers designed for it; a non-restartable reader forces from-zero replay, which is only safe because writers are idempotent. - *Chunk size:* correctness is unaffected, but it tunes restart cost vs throughput. - *Same-datasource transaction:* strongest atomicity, but cross-system reconciliation often can't share a transaction, pushing you toward idempotency keys and eventual-consistency reasoning. **Failure-mode audit.** Walk each: crash before any commit (restart re-reads, no effect); crash after chunk commit but before metadata (idempotent upsert absorbs the replay); operator double-click on completed date (rejected by guard); scheduler double-fire while running (rejected as already-running); partial external post (dedupe marker/idempotency key prevents re-charge). If all five are safe, the design holds. **Governance.** Add monitoring on `BATCH_JOB_EXECUTION` statuses, alert on repeated FAILED for a date, and give operators a controlled 'restart date X' vs 'abandon and advance' runbook rather than ad-hoc parameter fiddling.

  • Why not just use RunIdIncrementer so operators can always re-run without an exception?
    Because an ever-advancing run.id makes each trigger a new JobInstance, so two runs could both post the same business date — the exact double-post we must prevent. The business-date key intentionally trades effortless re-runs for the framework-level guarantee that a completed date can't run twice.
  • The reconciliation reads from one datasource and posts to a separate ledger system that can't join the batch transaction. How do you still avoid double-posting?
    Use an idempotency key the ledger honors (so it dedupes replays), or keep a local 'posted' table keyed by entry key that you write atomically before/with the post and check on replay. Without a shared transaction you rely on downstream idempotency plus your dedupe store, and design for at-least-once.
  • An operator genuinely needs to re-run a date that already COMPLETED (bad source data was fixed). How do you handle it safely?
    Provide a controlled runbook: either roll back/void the prior postings and start a fresh JobInstance under a new run key, or introduce a corrective re-run parameter. Because writers are idempotent upserts keyed by entry, a deliberate re-post overwrites cleanly rather than duplicating — but this must be an explicit, audited action, not an accidental relaunch.

saying these in an interview costs you the question

  • Slapping RunIdIncrementer on a job that must not double-post a given period
  • Assuming restart alone prevents duplicate side effects without idempotent writers
  • Treating cross-system posts as if they share the batch transaction
  • Letting operators re-run completed periods via ad-hoc parameter changes with no dedupe

context