How do you apply the State pattern to a long-lived business entity whose current state must be stored durably and whose transitions trigger external side effects such as sending an email or charging a card?
answer
- store stable code, not class name
- optimistic lock = compare-and-set on state
- decide purely, commit, then dispatch effects
- outbox + idempotency key = at-least-once safety
- durable timers, transition history, versioned definitions
basics
~20 sStore a stable code for the state, not the object, and rebuild the state object when loading. Decide transitions in memory, commit the new state, then perform outside effects afterwards — retried safely so a repeat does not charge twice.
solid answer
~60 sPersist a **stable discriminator** (`"IN_REVIEW"`), decoupled from class names, and reconstruct the state object through a lookup on load; unknown codes must fail loudly rather than default. Guard concurrency with optimistic locking (`UPDATE ... WHERE state = 'X' AND version = n`) or per-entity serialization — two events reading the same state and both transitioning is the classic corruption. Keep transitions **pure decisions**: compute the next state and a list of intended effects, commit the state change, then dispatch effects — never call a payment API inside the transition, since a rollback cannot un-charge a card and a commit failure after the call double-charges. Because a durable state change and an external call cannot share a transaction, use the **transactional outbox**: write the effect to the same database transaction and let a relay deliver it at-least-once, with idempotency keys derived from `(entityId, fromState, toState, attempt)`. Append a transition-history record for audit, make redelivered events that target the current state no-ops, and version the machine definition so in-flight entities are not stranded.
code
pseudocode · 19 lines// pure decision — trivially unit-testable, no I/O
AwaitingPayment.handle(PaymentCaptured e, order):
return Decision(next = Paid,
effects = [SendReceipt(order.email), NotifyWarehouse(order.id)])
// dispatcher: one place owns locking, atomicity, delivery
handle(orderId, event):
order = load(orderId)
d = order.state.handle(event, order)
if d.rejected: return
transaction {
rows = UPDATE orders SET state=d.next, version=version+1
WHERE id=? AND version=? // lost-update protection
if rows == 0: retry-or-drop // someone else moved it
INSERT transition_history(...)
for e in d.effects: // transactional outbox
INSERT outbox(payload=e, idempotency_key=hash(orderId, from, to, event.id))
}
// relay delivers outbox rows at-least-once; receivers dedupe on the keygo deeper
Say you store a code for the state and rebuild the state object when loading, and that emails or charges should happen after the state change is saved.
Add optimistic locking for concurrent events, a transition-history record for audit, and idempotency so a retried effect does not fire twice.
Describe pure decide-then-commit-then-dispatch, the transactional outbox with idempotency keys, no-op handling of redelivered events, and durable timers for deadline transitions.
Cover compensation and sagas for non-idempotent effects, event sourcing as an alternative source of truth, definition versioning and rolling-deploy compatibility, and the criteria for adopting a durable workflow engine instead of hand-rolling.
## What changes when state outlives the process In-memory, the State pattern is simple: swap a reference. For a database-backed entity three hard realities appear. 1. **You cannot store an object.** A row holds a string or enum, so you need a **bidirectional mapping** between the persisted discriminator and the state instance. 2. **Many actors touch the same entity.** Two threads, two pods, or a user click racing a webhook can both read state `X` and both decide to leave it. 3. **Transitions cause effects the database cannot roll back.** Emails sent, cards charged, warehouse messages published. A database transaction and an external API call are not atomic together. ## 1. Persisting the state ``` ORDER: id | state_code | version | updated_at ``` - **Stable codes, not class names.** Store `"AWAITING_PAYMENT"`; never derive it from a class name, or a refactor silently invalidates historical rows. Map codes → state objects through an explicit registry. - **Unknown codes fail loudly.** A row written by a newer deployment must not silently fall through to a default state during a rolling deploy; throw and let the entity be handled by an up-to-date instance. - **Removing a state is a migration**, not a code deletion: rows already sitting in it must be moved first, and old audit rows must still be interpretable. - **Keep a transition history table** — `(entity_id, from, to, event, actor, at, correlation_id)`. Auditors and support ask "who approved this and when"; reconstructing it later from logs is misery. If you go further and make the *event log* the source of truth, you are doing **event sourcing**: the current state becomes a fold over past events, giving perfect history and replay at the cost of schema-evolution discipline for old events. ## 2. Concurrency The canonical bug: two events read `AwaitingPayment`, both decide, both write. One transition is lost, or an illegal sequence executes (shipped *and* cancelled). Remedies, roughly in order of preference: - **Optimistic locking / conditional update.** `UPDATE orders SET state='PAID', version=version+1 WHERE id=? AND version=?`. Zero rows updated means someone beat you; reload and re-decide, or drop the event if it is now a no-op. This is a **compare-and-set on the state**, which is exactly what an FSM transition is. - **State-conditioned update** — `WHERE state='AWAITING_PAYMENT'` — same effect, and it reads as the transition's precondition. - **Per-entity serialization**: route all events for an entity id to one consumer/partition/actor. Removes the race entirely and is common in message-driven systems. - **Pessimistic row locks** (`SELECT ... FOR UPDATE`). Simple, but holds a lock for the duration and interacts badly with any external call made while holding it — another reason not to make such calls inside the transition. ## 3. Side effects — the core discipline **Split deciding from doing.** ``` result = state.handle(event, entity) // pure: no I/O, no clock, no randomness // result = (nextState, effects[]) persist(entity.withState(result.nextState), expectedVersion) // transactional dispatch(result.effects) // outside the transaction ``` Why this ordering: - **Effect inside the transaction, transaction later rolls back** → you charged a card for an order that does not exist. Databases cannot un-send. - **Effect before commit, commit then fails** → duplicate charge on retry. - **Commit first, then effect, and the process dies in between** → the effect is *lost*. This is the case people forget, and it is why fire-and-forget after commit is not good enough. The standard fix is the **transactional outbox**: in the *same* transaction that writes the new state, insert a row describing the intended effect. A separate relay reads unsent rows and performs them, marking them done. Now state change and effect intent commit atomically; delivery is **at-least-once**, so effects must be **idempotent**. **Idempotency in practice**: derive a deterministic key such as `hash(entityId, fromState, toState, eventId)` and pass it to the downstream API (most payment providers accept an idempotency key) or keep a processed-keys table. Duplicate delivery then becomes a no-op rather than a second charge. **Effects that cannot be made idempotent** (a physical action, a third-party call with no key) require **compensation**: model the failure as its own event that drives the machine to a compensating state (`PaymentFailed → AwaitingPayment`, `ShipmentRejected → Restocking`). This is the saga idea — no distributed transaction, just a state machine that knows how to undo. ## 4. Idempotent event handling Message infrastructure redelivers. If a `PaymentCaptured` event arrives twice and the entity is already `Paid`, the second must be a **no-op**, not an error and not a second transition. Two mechanisms: - **State-conditioned transitions**: the row-count-zero result of the conditional update tells you it already happened. - **Processed-event table** keyed by event id, written in the same transaction as the state change. Also decide **self-transition semantics** — does `Active --renew--> Active` re-run entry actions (restarting a timer) or not? Silent divergence here produces very confusing bugs. ## 5. Timers and deadlines "Cancel if unpaid after 7 days" is a transition triggered by *time*, and in-memory timers die with the process. Persist the deadline (a `due_at` column or a scheduled message) and have a scheduler emit a `Timeout` event that goes through exactly the same dispatch path as any other event. Anything else re-invents durable timers badly — the point at which adopting a durable workflow engine usually wins. ## 6. Evolving the machine States and transitions change while thousands of entities are mid-flight. Practices that hold up: - **Add before removing**: deploy code that understands a new state before anything writes it (important during rolling deploys where old and new instances run together). - **Version the definition** if it is data-driven; pin in-flight entities to the version they started under and migrate deliberately. - **Backfill** entities out of a state before deleting it, and keep the old code readable by history queries. ## 7. Where the state classes still earn their keep With all of the above, the concrete state classes become small and highly testable: they take an event and entity data and return `(nextState, effects)` with **no I/O at all**. Every transition rule becomes a pure unit test with no database and no mocks for the payment gateway. The messy parts — locking, outbox, retries, idempotency — live once in the dispatcher, not once per state. That separation is the real reason to keep the pattern at this scale rather than sprinkling status checks through service methods.
- Why not simply send the email inside the transaction that updates the state?Because the two cannot be atomic. If the transaction later rolls back you have sent a message about something that did not happen; if you commit after the call and the commit fails, the retry sends it twice. Writing the intent to an outbox row in the same transaction makes the state change and the intent atomic, and delivery becomes a separate at-least-once concern.
- An event is redelivered and the entity is already in the target state. What should happen?A no-op, not an error. Detect it via a state-conditioned update that affects zero rows, or a processed-event table keyed by event id written in the same transaction. Treating redelivery as an illegal transition turns normal at-least-once messaging into a stream of false alarms.
- How do you implement 'cancel the order if it is unpaid after seven days'?Persist the deadline — a due_at column or a scheduled message — and have a scheduler emit a Timeout event through the same dispatch path as any other event. In-process timers do not survive restarts or rebalancing, so durable scheduling is required; needing many such rules is a strong signal to adopt a durable workflow engine.
- What breaks during a rolling deploy that introduces a new state?Old instances load a row whose state code they do not recognise. Deploy code that understands the new state before any code writes it, and make unknown codes fail loudly rather than falling back to a default — a silent default can move an entity backwards or skip guards.
A shipping container's status board. The board (database) records where the container is; the paperwork triggered by each move (customs filing, invoice) goes out after the move is officially recorded, stamped with a reference number so a re-sent copy is filed once, not twice. You never dispatch the paperwork first and hope the move sticks.
saying these in an interview costs you the question
- Persisting the state class name, so a rename or package move silently breaks historical rows.
- Calling an external API inside the transition or the transaction, making rollback impossible and retries double-charge.
- Assuming exactly-once delivery — outbox relays and message brokers are at-least-once, so effects must carry idempotency keys.
- Handling a redelivered event as an illegal transition instead of a no-op, producing constant false alarms in an at-least-once system.
- Relying on in-process timers for deadline transitions, which vanish on restart or rebalance.
- Skipping optimistic locking and assuming a single writer, then losing transitions the first time a webhook races a user action.
- Deleting a state from the code while rows are still sitting in it, stranding entities on the next load.