Your team is deciding whether to build a stateful system on the actor model rather than stateless services over a shared database. What do actors give you, where does the model break down, and how would you decide?
answer
- fits: many independent entities, hot mutable state
- per-entity serialization replaces locking/retries
- breaks: cross-entity transactions, hot keys, lost stack traces
- durability is opt-in: event sourcing + snapshots
- hybrid: actors for the hot core, database for the rest
basics
~20 sActors fit entity-shaped, stateful, high-touch domains: isolated state, serialized per-entity access, supervision, natural sharding. They break down on cross-entity transactions, hot keys (one actor equals one core), lost stack traces and harder debugging, weak delivery guarantees, and state that must survive restarts. Decide by whether per-entity ownership is the dominant access pattern.
solid answer
~60 s**Where actors win:** the domain is a large population of independent entities with identity and mutable state accessed repeatedly — devices, sessions, game objects, carts, connections. Per-entity serialization removes the read-modify-write races that make the database path need locking or optimistic retries; keeping state in memory removes a round trip per operation; supervision gives per-entity fault isolation; addressing gives you a sharding key for free. **Where it breaks down:** anything spanning entities needs an explicit protocol (sagas, two-phase commit) because there is no cross-actor transaction; a hot entity is a single-core bottleneck since processing is serial; asynchronous messaging destroys stack traces and makes debugging a tracing exercise; delivery is at-most-once with only pairwise ordering; state must be made durable (event sourcing/snapshots) or a restart loses it; and location transparency can hide network failure and latency. **Decide** by asking whether per-entity ownership dominates, whether hot keys can be sharded, whether the team can operate a stateful cluster, and whether the durability and distributed-tracing work is affordable. If interactions are mostly cross-entity queries and transactions, stateless services over a database are simpler and better.
go deeper
Recognize that actors suit stateful entity-shaped problems and that they do not provide database-style transactions across entities.
Contrast the two designs concretely — per-entity serialization versus locking or optimistic retries — and name the main costs: durability, debugging, weaker delivery guarantees.
Drive the breakdown list with mitigations: sagas for cross-entity invariants, sharding for hot keys, correlation ids and tracing, event sourcing for durability, and the realities of operating a stateful cluster.
Make the call explicitly against the dominant access pattern, propose the hybrid boundary, weigh team operational capability, and state the reversal cost of each option so the decision is defensible later.
## Frame the choice honestly The actor model is not a general replacement for stateless services; it is a strong fit for one shape of problem and an awkward fit for others. The decision is about the *dominant access pattern*, plus the operational cost the team can carry. ## What the model genuinely gives you **Per-entity serialization for free.** One actor per entity means all operations on that entity are processed one at a time. The read-modify-write race that forces pessimistic locking or optimistic-retry loops in a database design simply does not exist. For high-contention entities this is a real throughput win, not just a simplicity win. **In-memory state with an owner.** The entity's state lives where the work happens, so a hot entity does not pay a database round trip per operation. Because exactly one actor owns it, there is no cache-coherence problem across nodes. **Fault isolation with a policy.** Supervision means a failure damages one entity's state and a parent decides the response, rather than an exception propagating through a shared request path. **Natural sharding and location transparency.** Addressing by identity gives you the partition key, and messages can cross a network without changing the programming model. **Domain fit.** Anything with identity and a lifecycle — a device, a session, a match, a subscription — maps to an actor almost one-to-one, and state machines are pleasant to write as behaviour changes. ## Where it breaks down **Cross-entity consistency.** There is no transaction spanning actors. Any invariant over two entities (transfer funds, reserve inventory across warehouses) needs an explicit protocol: a saga with compensations, a coordinator actor, or a designated owner for the combined invariant. Teams used to a database transaction routinely underestimate this, and it is the single most common reason an actor design becomes painful. **Hot keys.** Serialization is per actor, so one entity's throughput is capped at one core. A celebrity account, a popular product, a global counter — each becomes a bottleneck. The fix is to split it (per-shard sub-actors that aggregate) which reintroduces cross-entity coordination. **Debuggability.** Asynchronous messaging discards the call stack. 'Why did this actor receive this message' requires correlation ids and distributed tracing that you must build in from day one; retrofitting it is expensive. Nondeterministic interleaving also makes tests flakier and reproduction harder. **Weak guarantees by default.** At-most-once delivery, pairwise-only ordering, and no automatic retry mean protocols need timeouts, idempotency and de-duplication in code that would be implicit in a synchronous call. **Durability is your problem.** In-memory state vanishes on restart or rebalance. You must add event sourcing with snapshots or write-through persistence — a substantial subsystem with its own schema-evolution and replay concerns. **Cluster operations.** A stateful cluster brings membership, split-brain protection, rebalancing when nodes join or leave, and rolling upgrades that must move live state. Stateless services over a database push all of that onto the database, which is usually the more familiar operational burden. **Typing and protocol discipline.** In many runtimes messages are loosely typed, so protocol errors surface at runtime; keeping a protocol coherent across many actor types requires deliberate design and review. ## How I would decide Approve actors when most of these hold: a large population of independent entities; repeated stateful interactions per entity (not a single read); genuine contention per entity that a database would have to lock around; latency requirements that make a database round trip per operation costly; a natural partition key with no unfixable hot key; a team willing to own durability, tracing and cluster operations. Prefer stateless services over a database when: the workload is dominated by cross-entity queries, reports and multi-entity transactions; state is naturally relational and consistency requirements are strong; the team's operational strength is the database; or the traffic simply does not justify keeping state in memory. **Hybrids are usually right.** Use actors for the hot, stateful, entity-shaped core and keep everything else stateless with a database as the source of truth. Also consider the middle options: partitioned stateless consumers over a log (per-key ordering with durability, no in-memory ownership), or channel/CSP-style pipelines when the shape is a data flow rather than a population of entities. Finally, name the reversal cost. Moving an entity's state from an actor into a database later is a moderate refactor if durability was already event-sourced, and a rewrite if state exists only in memory. That argues for persisting from the start even when it is not yet needed. ## Interview delivery Give the fit criteria, the four or five real breakdowns (cross-entity transactions, hot keys, debuggability, weak delivery guarantees, durability and cluster ops), state the decision rule in terms of the dominant access pattern, and volunteer the hybrid as the usual answer. Concluding with the reversal cost signals genuine decision-making rather than technology preference.
- How do you handle an operation that must atomically update two entities owned by different actors?There is no cross-actor transaction, so you make the protocol explicit: either one actor becomes the owner of the combined invariant and both updates happen inside it, or you run a saga — perform the steps in sequence with compensating actions if a later step fails, and make each step idempotent because messages may be retried. Both approaches trade atomicity for eventual consistency, so the domain has to tolerate a visible intermediate state.
- One entity receives a hundred times the traffic of the others. What do you do?Split it, because a single actor's throughput is capped at one core by serialized processing. Typically you shard the entity into N sub-actors keyed by a secondary dimension and aggregate their results, or separate reads from writes by serving reads from a replicated snapshot. Both reintroduce coordination, which is exactly the cost of removing the bottleneck.
- What is the first thing you would build into an actor system that teams commonly retrofit too late?Correlation identifiers and distributed tracing on every message, because asynchronous messaging destroys the call stack and without them production debugging is guesswork. A close second is durability of actor state — event sourcing or snapshots — since retrofitting persistence onto state that only ever existed in memory is close to a rewrite.
saying these in an interview costs you the question
- Presenting actors as a general replacement for stateless services and databases
- Assuming transactions work across actors
- Ignoring hot keys given that per-actor processing is serial
- Treating location transparency as if remote sends have the same failure profile as local ones
- Planning to add persistence and tracing 'later'