skip to content

questions

6

What is an event store, and how does its append-only log differ from a traditional database table that you update in place?

level: juniorimportance: must knowfreq 55%

answer

  1. append-only, no UPDATE/DELETE
  2. events are facts, past tense
  3. current state = fold over stream
  4. compensating events not edits
  5. snapshot to bound replay cost

basics

~20 s

An event store only lets you add new records, never change or delete old ones. Instead of overwriting a row to reflect current state, it keeps every change as its own permanent entry, in order.

solid answer

~40 s

An event store is a database specialized for persisting a sequence of immutable events instead of mutable rows. Where a normal CRUD table represents only current state (UPDATE overwrites history), an event store appends every state-changing fact as a new record and never edits or deletes existing ones. The current state of an entity is derived by replaying its events in order, not stored directly (though snapshots may cache it). This log is the system of record: if you need to know why the current balance is $400, you don't inspect a snapshot, you replay the events that produced it. This gives you built-in audit trail, time-travel debugging, and the ability to build new read models later from history you already have, at the cost of write-side complexity and read-time reconstruction cost.

go deeper

for a junior

Should describe append-only in own words and know that current state comes from replaying events, without needing to discuss snapshotting or GDPR trade-offs.

for a middle

Should mention snapshots as the fix for replay cost and know compensating events are how corrections happen, not edits.

for a senior

Should discuss schema evolution/upcasting, storage growth trade-offs, and be able to compare event store to CRUD table with concrete cost/benefit language.

for a principal

Should reason about when NOT to use event sourcing/append-only stores at all — e.g., simple CRUD domains where audit trail has no business value — and discuss GDPR erasure strategy and cross-team schema versioning discipline.

## The single invariant An event store is a persistence mechanism built around a single invariant: once a record (an event) is written, it is never modified or deleted. Mechanically, writing to an event store means calling an append operation that adds one or more events to the end of a named stream — a stream is typically the ordered history for one entity, such as `account-482` or `order-91a3`. - Each event carries a payload describing something that already happened — `MoneyDeposited`, `OrderShipped`, `EmailChanged` — expressed in the **past tense**, plus metadata: a stream name, a position/sequence number within that stream, a timestamp, and often a causation/correlation ID. - There is no `UPDATE` or `DELETE` API surface exposed to application code; if a mistake needs correcting, you append a new **compensating event** (`DepositReversed`) rather than editing the wrong one out of history. ## How it differs from a table you update in place This differs fundamentally from how a relational table backing a typical CRUD service works. An accounts table with a balance column represents only the current truth: an `UPDATE accounts SET balance = 400 WHERE id = 482` physically overwrites the previous value, and unless you've bolted on a separate `audit_log` table, the fact that the balance used to be 350, and why it changed, is gone. The event store makes **the sequence of changes** the primary artifact and **the current state a derived value**: to know an account's balance you fetch the stream `account-482`, get back `[AccountOpened, Deposited:100, Deposited:250, Withdrawn:50, ...]`, and fold over them in order to arrive at 400. This process is called **replaying** (or 'hydrating') the aggregate, and it's the same mechanism used whether you're: - rebuilding in-memory state to handle the next command, - debugging what happened last Tuesday, - or spinning up a brand-new read model six months after the events were written. ## Why the model exists The reason this model exists is that in many domains the history of what happened is itself valuable business data, not incidental exhaust. - **A banking system, an inventory system, or a shipping system** often needs to answer 'why is this the current state' for compliance, dispute resolution, or analytics — and a mutable table structurally throws that answer away on every write. - **Event sourcing also decouples the write model from the read model**: because the log is immutable and ordered, you can build arbitrarily many projections (current-balance view, monthly-statement view, fraud-detection view) from the same underlying facts, including projections you didn't anticipate when the system was first built, simply by replaying the log from the beginning against new folding logic. ## The trade-offs The trade-offs are real on both sides. On the cost side: - reconstructing current state means replaying potentially thousands of events per aggregate unless you introduce **snapshots** (periodic materialized checkpoints) to bound replay cost — so raw event-store reads are slower and more CPU-intensive than a single indexed row lookup. - Storage grows monotonically since nothing is ever deleted, which complicates 'right to be forgotten' requirements (GDPR erasure typically requires either **crypto-shredding** — encrypting PII per-subject and discarding the key — or a documented exception, because you literally cannot `DELETE` a row). - Schema evolution is harder too: old events on disk were serialized against yesterday's event shape, so your read side must tolerate multiple versions of the same event type forever (**upcasting**), rather than relying on a single current schema like a normal table would. On the benefit side, you get: - a genuine audit log for free, - the ability to answer new business questions retroactively by replaying history through new logic, - and natural support for temporal queries ('what did this order look like at 3pm'). ## Failure modes Failure modes tend to surface at scale or under operational pressure. 1. A common one is **unbounded replay cost**: teams skip snapshotting early on, and an aggregate that accumulates tens of thousands of events (a long-lived shopping cart, a chatty IoT device stream) starts timing out on load because every command handler replays the full history before it can validate the next command. 2. Another is **treating the event store like a queue** and letting consumers 'delete after processing' — this defeats the entire point and quietly turns the append-only log into a lossy buffer, so any later replay or new projection is working from an incomplete history. 3. A third is silent **event-shape drift**: a producer changes a payload field's type or renames it without an upcaster, and every consumer that deserializes older events starts throwing or silently mis-parsing. ## Where it shows up - The canonical example is **EventStoreDB** (formerly Event Store), purpose-built around this append-only-stream model with per-stream optimistic concurrency. - **Apache Kafka** with infinite retention plus compaction disabled is often pressed into a similar role, though it lacks per-key consistency guarantees an event store gives natively. Domain-driven design communities most commonly cite banking ledgers and e-commerce order pipelines as the textbook use case, precisely because 'what happened and in what order' is the actual business requirement, not an afterthought.

  • If a mistake gets appended to the stream, like a duplicate deposit event, how do you fix it without breaking immutability?
    You append a compensating event, such as DepositReversed or DepositCorrected, that cancels or adjusts the effect of the mistaken one when folded. The wrong event stays in history forever, but the derived current state (and audit trail) correctly reflects both the mistake and its correction, which is actually valuable for auditors.
  • How does an event store avoid replaying thousands of events every time an aggregate is loaded?
    Most implementations periodically write a snapshot — a materialized copy of the folded state at a given stream position — so loading only needs to fetch the snapshot plus the (small number of) events appended after it. The snapshot is a cache, not a source of truth; it can always be rebuilt by replaying from position zero.
  • How do you handle GDPR 'right to be forgotten' if you can never delete an event?
    The common technique is crypto-shredding: personally identifiable fields are encrypted per-subject with a dedicated key, and 'deleting' the data means discarding that key so the ciphertext left in the immutable event becomes permanently unreadable, without violating the append-only guarantee.

Like a bank passbook/ledger vs. a whiteboard: the ledger only ever gets new lines added at the bottom (deposit, withdrawal), and your current balance is what you get by adding them all up — you never erase an old line to 'fix' the balance, you write a correcting entry.

saying these in an interview costs you the question

  • says you 'just UPDATE the event' to fix a mistake
  • doesn't know current state is derived by replay
  • thinks event stores don't need snapshots at any scale
  • conflates event store with a generic message queue with no persistence guarantee
  • unaware immutability complicates GDPR erasure

context

open as a page

When two commands try to append to the same event stream at nearly the same time, how does optimistic concurrency control in an event store prevent one write from silently clobbering the other?

level: middleimportance: must knowfreq 75%

basics

~20 s

Each write says 'I'm adding this assuming the stream currently has N events.' If someone else already added an event in between, the store rejects the write instead of accepting it blindly, so the app can retry with fresh data.

open as a page

Why do event-sourced systems typically partition the event log into one stream per aggregate ID rather than writing every event for every entity into a single global stream?

level: middleimportance: must knowfreq 65%

basics

~20 s

Grouping by aggregate id keeps each entity's own history together and in order, so you can load just that one entity's events fast, and two different entities' writes never block or interleave with each other.

open as a page

When a projection subscribes to read events from an event store to build a read model, what ordering guarantees can it actually rely on, and where do they break down?

level: seniorimportance: must knowfreq 60%

basics

~20 s

Within one entity's own history, events always arrive in the exact order they happened. But if you're watching many entities at once, events from different ones can arrive interleaved in almost any order relative to each other, so you can't assume a global timeline unless the store gives you one explicitly.

open as a page

A client retries an append call to an event store after a network timeout, but the original append had actually already succeeded server-side before the timeout. How do idempotency keys prevent this from creating a duplicate event?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Each write carries a unique ID. If the same ID shows up twice because of a retry, the store recognizes it already processed that exact write and just confirms success again instead of adding a second copy.

open as a page

You're designing an event-sourced system where a small number of aggregates — say, the top few 'hot' inventory SKUs during a flash sale — receive a disproportionate share of all writes, causing constant optimistic-concurrency retries on those specific streams. What are the real options for fixing this, and what does each one cost?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

When too many people try to update the same one record at once, retries pile up. Fixes usually mean either spreading that one record's updates across several smaller pieces, changing the kind of update so order stops mattering, or accepting some updates outside the strict record and reconciling later — each trades away some strict consistency or simplicity for more speed.

open as a page