When would you actually adopt Spring Statemachine in production, and how do you persist a machine's state across requests or restarts?
answer
- Use when lifecycle rules must be enforced; skip for toggles
- Machine = stateful, not thread-safe
- Factory + rehydrate per request; DB is source of truth
- StateMachineContext = state(s)+extended vars
- StateMachinePersist / StateMachinePersister (Redis/JDBC/Mongo)
basics
~20 sUse it when a domain has a real, rule-bound lifecycle where illegal transitions must be prevented; skip it for trivial toggles. Persist by snapshotting the machine's state into a StateMachineContext (via StateMachinePersister) and restoring it per request/aggregate.
solid answer
~50 sAdopt Spring Statemachine when the lifecycle is genuinely stateful and rule-heavy — many states, guarded transitions, entry/exit behaviour, and a hard requirement that illegal transitions can't happen (order/payment/approval/device lifecycles). For a two-state flag it's overkill. Because a StateMachine<S,E> is a live, stateful, not-thread-safe object, in a stateless web app you don't keep one per user in memory; you use @EnableStateMachineFactory to build a fresh machine per request and rehydrate its state. Persistence uses StateMachinePersister<S,E,T> plus a StateMachinePersist implementation that serialises a StateMachineContext (current state(s), extended-state variables, region info) to a store — JDBC, Redis (spring-statemachine-redis), or MongoDB starters exist. On each request you persist().save the context after handling and restore() before, keyed by the business id. Key concerns: the machine isn't thread-safe (guard the per-aggregate instance), context must capture all region states, and version/evolve the state enum carefully since old persisted contexts must still deserialise.
code
java · 25 lines// Persist-per-request pattern with a factory + persister.
@Service
public class OrderWorkflow {
private final StateMachineFactory<OrderState, OrderEvent> factory;
private final StateMachinePersister<OrderState, OrderEvent, String> persister;
public OrderWorkflow(StateMachineFactory<OrderState, OrderEvent> factory,
StateMachinePersister<OrderState, OrderEvent, String> persister) {
this.factory = factory;
this.persister = persister;
}
public void handle(String orderId, OrderEvent event) throws Exception {
StateMachine<OrderState, OrderEvent> machine = factory.getStateMachine(orderId);
persister.restore(machine, orderId); // rehydrate from store
machine.sendEvent(Mono.just(
MessageBuilder.withPayload(event).setHeader("orderId", orderId).build()
)).blockLast();
persister.persist(machine, orderId); // snapshot new state back
}
}
// StateMachinePersister wraps a StateMachinePersist<OrderState,OrderEvent,String>
// (e.g. RepositoryStateMachinePersist / Redis / JDBC) that serialises the
// StateMachineContext for orderId.go deeper
Can say state can be saved to a database and reloaded; may not know the mechanism.
Knows to use a factory per aggregate and that some persister saves/restores state.
Explains StateMachineContext + StateMachinePersist/Persister, the not-thread-safe constraint, and rehydrate-per-request.
Makes the build/adopt call against alternatives, designs persistence + concurrency + enum-evolution strategy, and treats transitions as an auditable, enforceable spec.
## Is it the right tool? Spring Statemachine earns its keep when a domain has a **real lifecycle with enforceable rules**: - Many discrete states and a diagram that *is* the spec (order processing, payment capture, KYC/approval workflows, device/connection lifecycles, provisioning). - **Illegal transitions must be impossible** — the value is that the framework rejects an event with no valid guarded transition, centralising the rules. - Behaviour on **entry/exit**, **guards**, hierarchy, or parallelism matters. It's the *wrong* tool when: the lifecycle is a trivial boolean/two-state toggle (a field is cheaper); the 'workflow' is really long-running human-task orchestration with timers, escalation and audit (a BPM/workflow engine or a durable orchestrator like Temporal/Camunda fits better); or you need distributed saga coordination across services (a saga/orchestration pattern is the better frame, though a state machine can implement each local step). Teams also weigh the **added dependency and learning curve** against just writing explicit guarded service methods. ## The stateful-object problem A `StateMachine<S,E>` is a **live, mutable, not-thread-safe** object holding a current state. That collides with a **stateless, multi-threaded web tier**. Two viable models: 1. **Machine-per-aggregate, rehydrated per request (typical):** use `@EnableStateMachineFactory` to build a fresh machine each request, **restore** the persisted state for that business id, handle the event, **save** the new state, discard the machine. The database (not JVM memory) is the source of truth. 2. **Long-lived in-memory machine** only for genuinely singleton, single-threaded control (e.g. a device controller), with your own concurrency guarding. ## Persistence mechanics The machine's snapshot is a **`StateMachineContext<S,E>`** — it captures the **current state(s)** (all region states for parallel machines), the **extended state variables**, history, and machine id. Persistence is layered: - **`StateMachinePersist<S,E,T>`** — low-level interface: `write(context, contextObj)` / `read(contextObj)`. You implement how a `StateMachineContext` maps to/from your store (T is the context key, e.g. an order id). - **`StateMachinePersister<S,E,T>`** — higher-level helper wrapping a `StateMachinePersist`: `persister.persist(machine, id)` and `persister.restore(machine, id)`. `DefaultStateMachinePersister` is the common implementation. - **Ready-made stores:** starters exist for **Redis** (`spring-statemachine-redis`) and **MongoDB**, and JDBC-based persistence — so you don't hand-roll serialisation. Typical request flow: `restore(machine, businessId)` -> `sendEvent(...)` -> if accepted, `persist(machine, businessId)`. ## Production concerns and gotchas - **Thread-safety:** never share one machine instance across concurrent requests for the same aggregate without external locking (optimistic version on the row, or a distributed lock). Rehydrate-per-request sidesteps most of it but concurrent events on the *same* id still race — guard with the store's version. - **Capturing all region states:** for orthogonal regions the context must persist *every* active region state; a naive single-state save loses parallel state. - **Enum/schema evolution:** persisted contexts reference your state/event enum names. Renaming or removing an enum constant can break deserialisation of old rows — evolve additively and migrate carefully. - **Extended-state serialisation:** variables you stash must be serialisable to your store; keep them small and stable. - **Autostart vs restore:** when rehydrating you usually don't want autoStartup re-running the initial-state entry action; restore sets the state directly. - **Error handling:** action exceptions can push the machine into an error state — decide whether that's persisted or rolled back with the surrounding transaction. - **Observability:** register a `StateMachineListener` for transition logging/metrics; treat the transition log as an audit trail.
- Why not keep one StateMachine instance per user in an HTTP session?A machine is a stateful, non-thread-safe object; sticky per-session instances waste memory, don't survive restarts or scale-out, and race under concurrent requests. Better to make the datastore the source of truth and rehydrate a fresh machine per request via the factory + persister.
- What exactly does StateMachineContext capture, and why does it matter for parallel machines?It captures the current state(s), extended-state variables, history and machine id. For orthogonal regions there are multiple active states, so the context must record every region's state; persisting a single state would silently lose the parallel state on restore.
- What breaks when you rename a state enum constant after go-live?Persisted StateMachineContext rows reference enum names; renaming/removing a constant can fail to deserialise existing rows, effectively corrupting in-flight machines. Evolve the enum additively and migrate stored contexts rather than renaming in place.
- When would you pick a workflow/BPM engine over Spring Statemachine?For long-running, human-in-the-loop processes needing timers, escalation, task assignment, audit and visual authoring (Camunda/Flowable), or durable distributed orchestration (Temporal). Spring Statemachine fits in-process, code-defined lifecycles where enforcing legal transitions is the main goal.
saying these in an interview costs you the question
- Reaching for a state machine on a trivial two-state toggle
- Sharing one stateful machine across concurrent requests without locking
- Persisting only a single state for a parallel (multi-region) machine
- Renaming state enum constants in place after there are persisted contexts
- Assuming the machine survives restarts on its own without explicit persistence