How does incomplete-event recovery work in Spring Modulith, and how would you operate it in production (restart replay, scheduled resubmit, cleanup, idempotency)?
answer
- incomplete row = listener never completed
- republish-outstanding-events-on-restart (default false)
- IncompleteEventPublications.resubmitIncompletePublicationsOlderThan(Duration)
- CompletedEventPublications.deletePublicationsOlderThan for cleanup
- no built-in retry cap → handle poison events yourself
basics
~10 sFailed listeners leave incomplete publication rows. Recover them by enabling republish-outstanding-events-on-restart, or by scheduling a job that calls IncompleteEventPublications.resubmitIncompletePublicationsOlderThan(...). Purge completed rows with CompletedEventPublications. Because resubmits re-run listeners, handlers must be idempotent.
solid answer
~40 sAn incomplete publication is a registry row whose listener never completed — from a thrown exception or a crash before completion. Two recovery paths: at startup, spring.modulith.events.republish-outstanding-events-on-restart=true resubmits everything still incomplete; and at runtime you inject the IncompleteEventPublications bean and call resubmitIncompletePublicationsOlderThan(Duration) or resubmitIncompletePublications(Predicate), typically from a @Scheduled job so transient failures self-heal without a restart. Filtering by age avoids racing an in-flight listener. Because resubmission re-invokes the listener, delivery is at-least-once and handlers must be idempotent. For hygiene, choose a completion-mode (UPDATE keeps rows, DELETE/ARCHIVE keep the live table small) and periodically purge completed rows via CompletedEventPublications.deletePublicationsOlderThan. Watch for poison events that fail forever — add alerting on old incomplete rows and consider a dead-letter or max-attempt strategy, since Modulith itself has no built-in backoff cap.
go deeper
Know that failed events can be re-sent on restart.
Name republish-on-restart and that IncompleteEventPublications can resubmit programmatically.
Design a scheduled resubmit with an age filter, cleanup via completion-mode/CompletedEventPublications, and idempotent handlers.
Own the operational story: poison-event quarantine, metrics/alerting on backlog, multi-instance replay coordination, and event schema evolution policy.
## What 'incomplete' means Every transactional listener gets an `EVENT_PUBLICATION` row when the event is published. The row is **incomplete** until that specific listener finishes successfully (no `COMPLETION_DATE`, or still present depending on completion mode). It becomes incomplete-and-stuck when: - the listener **threw** (async exception, logged, not retried automatically at that moment), or - the JVM **crashed** after commit but before the listener completed, or - the listener **succeeded** but the process died before the completion was written (this is why replay can double-fire → at-least-once). ## Recovery mechanisms ### 1. Republish on restart ```properties spring.modulith.events.republish-outstanding-events-on-restart=true ``` Default is **false**. When true, on application startup Spring Modulith reads all still-incomplete publications and **resubmits** them to their listeners. Good for surviving deploys/crashes, but it only fires at boot — a long-running app that hits a transient failure won't retry until it restarts. Also be mindful in **multi-instance** deployments: you want a single owner replaying, not every instance re-firing the same rows (coordinate via configuration/leadership or run replay only where appropriate). ### 2. Programmatic / scheduled resubmit Inject the bean and drive retries yourself: ```java @Component requiredArgs class EventRetryJob { private final IncompleteEventPublications incomplete; @Scheduled(fixedRate = 60_000) void retry() { // only things stuck longer than 5 minutes, so we don't race in-flight listeners incomplete.resubmitIncompletePublicationsOlderThan(Duration.ofMinutes(5)); } } ``` Or filter precisely with `resubmitIncompletePublications(Predicate<Object> filter)` where the predicate receives the deserialized event, letting you resubmit only certain types. **Why the age filter matters:** without `OlderThan`, a scheduled resubmit could grab a publication whose listener is *currently running*, causing a needless concurrent re-invocation. The duration acts as a simple 'give it time to finish' guard. ### 3. Cleanup of completed publications If you use `completion-mode=UPDATE` (the default), completed rows accumulate. Purge them: ```java completed.deletePublicationsOlderThan(Duration.ofDays(7)); ``` via the `CompletedEventPublications` bean, or switch `spring.modulith.events.completion-mode` to `DELETE` (row removed on completion) or `ARCHIVE` (moved to an archive table) so the hot table stays small — important because incomplete-scanning queries run against it. ## Operating it well (the senior/principal part) - **Idempotency is mandatory.** Every recovery path re-invokes the listener; plus the crash-after-side-effect case guarantees genuine duplicates. Use idempotency keys / upserts / processed-markers. - **Poison events.** Modulith resubmits but has **no built-in max-attempt/backoff cap**. An event that always fails will retry forever. Add: alerting on publications older than N minutes/hours, a metric on incomplete count, and an application-level dead-letter or attempt counter to quarantine poison events. - **Observability.** Expose the incomplete count as a gauge; a rising backlog signals a downstream outage. Spring Modulith surfaces publication info and there's actuator/insight support for inspecting the registry. - **Schema evolution.** Incomplete events sit serialized in the table across deploys. If you rename fields/types, an old payload may fail to deserialize on replay. Keep events additive/backward-compatible, or migrate stored rows. - **Multi-instance & ordering.** Replays don't preserve global ordering; design handlers that don't assume it. ## Summary Recovery = durable incomplete rows + (restart replay | scheduled resubmit) + cleanup, all resting on idempotent, evolution-tolerant handlers, with your own guardrails for poison events since the framework won't cap retries.
- Why pass a Duration to resubmitIncompletePublicationsOlderThan instead of resubmitting everything?To avoid racing listeners that are currently in-flight. The age threshold means 'this has been stuck long enough that it's genuinely failed, not merely still running,' preventing redundant concurrent re-invocations from a scheduled job.
- Spring Modulith retries incomplete events but doesn't cap attempts. How do you handle a poison event?Add your own guardrails: an attempt counter or age-based quarantine, an application-level dead-letter table, and alerting/metrics on incomplete count and oldest incomplete age. The framework won't stop retrying a permanently failing event on its own.
- What risk do stored serialized events pose across a deployment?Schema drift. An incomplete event serialized under the old schema must still deserialize under the new code on replay. Keep event changes backward-compatible (additive fields, tolerant readers) or migrate the stored rows, or the replay will fail.
saying these in an interview costs you the question
- Assuming republish-on-restart is enabled by default (it's false)
- Resubmitting all incompletes without an age filter and racing in-flight listeners
- Believing Modulith caps retries / dead-letters poison events automatically
- Ignoring table growth when completion-mode is UPDATE
- Not accounting for event schema evolution on replay