skip to content

Event Sourcing

Storing the sequence of things that happened and deriving current state by replaying it, instead of storing the state itself. Interviewers probe it because it changes what fixing a bug means.

part ofEvent-driven architecture & messagingoverview, primer and where to startread it →
on this pageshow

questions

page 2 of 2

You're designing an event-sourced system where a small number of aggregates — say, the top few 'hot' inventory SKUs during a flash sale — receive a disproportionate share of all writes, causing constant optimistic-concurrency retries on those specific streams. What are the real options for fixing this, and what does each one cost?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

When too many people try to update the same one record at once, retries pile up. Fixes usually mean either spreading that one record's updates across several smaller pieces, changing the kind of update so order stops mattering, or accepting some updates outside the strict record and reconciling later — each trades away some strict consistency or simplicity for more speed.

open as a page

In a system that's been through a dozen event schema revisions over many years, how do deep upcaster chains become an operational problem, and what strategies let a team retire old versions rather than carrying every translation step forever?

level: principalimportance: nice to knowfreq 20%

basics

~20 s

If an event type has been changed a dozen times, the oldest events might need to run through eleven small translation steps every time they're replayed, which is slow and hard to maintain. Teams fix this by periodically rewriting old streams into the current shape once, so future reads skip the whole chain.

open as a page

A projection backing a high-traffic, customer-facing read API takes four hours to fully rebuild from the event log after a handler bug fix. What techniques let you deploy the corrected projection without taking that read API offline during the rebuild?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Build the new, corrected version in a completely separate table while the old one keeps serving traffic, and only switch customers over to the new one once it's fully caught up and checked - like renovating a store in a new building instead of closing the old one down for the day.

open as a page

showing 31–33 of 33