Redis Enterprise offers active-active geo-replication, where several regions accept writes to the same dataset simultaneously. What mechanism makes concurrent writes converge without a coordinator, and which guarantees do you NOT get from it?
answer
- CRDT merge: commutative, associative, idempotent — no coordinator
- counters replicate deltas → increments sum
- collections add-wins (observed remove); strings last-write-wins
- TTL converges to the longer; deletes leave tombstones
- no linearizability, no cross-region WATCH/MULTI, async = possible loss
basics
~20 sEach region holds a full replica built on conflict-free replicated data types. Writes are applied locally, then replicated asynchronously and merged by per-type rules: counters sum, sets and hashes are add-wins, plain string values resolve last-write-wins. You do not get global linearizability, cross-region atomicity, or loss-free failover.
solid answer
~60 sActive-active databases (CRDBs) give every participating region a full, writable copy. Writes are accepted locally with local latency, then shipped asynchronously to the other regions, where merge rules — **CRDT semantics baked into each data type** — resolve concurrency deterministically without any coordinator or quorum: - **Counters** (INCR/DECR, hash-field and sorted-set increments) are operation-based: each region replicates the *delta*, so concurrent increments **sum** rather than overwrite. - **Sets, hashes, sorted-set membership** are observed-remove/add-wins: a concurrent add and delete of the same element resolves to the element existing. - **Plain string values** have no mergeable structure, so concurrent SETs resolve **last-write-wins** by a clock the system maintains; one of the writes is discarded. - **Expiry** converges by taking the longer TTL, and deletes leave tombstones until garbage collected. What you do **not** get: linearizability or global ordering; read-your-writes across regions; cross-region transactions (MULTI/EXEC and WATCH are local only); a guarantee that an acknowledged local write survives losing that region before it replicates. Memory overhead is also higher because of CRDT metadata and tombstones.
go deeper
Know that every region can accept writes and that conflicts are resolved automatically by per-data-type rules rather than by a coordinator.
Distinguish the merge behaviours: counters sum, collections are add-wins, plain strings are last-write-wins.
State the guarantees you lose — no linearizability, no cross-region WATCH/MULTI, possible loss on regional failure — and design around them with deltas and region affinity.
Treat it as a data-modelling commitment: pick active-active only when the workload is dominated by mergeable operations, and price the full-copy-per-region storage and write fan-out.
## The problem being solved A globally distributed application wants local write latency in every region. The classical alternatives are a single primary (remote regions pay a round trip and lose write availability when the primary is unreachable) or consensus (every write pays a quorum round trip). Active-active takes a third route: **accept writes everywhere, replicate asynchronously, and make conflicts impossible to get wrong by construction**. "By construction" means conflict-free replicated data types: data types whose merge function is commutative, associative and idempotent, so replicas that have seen the same set of updates — in any order, with duplicates — reach the same state. No coordinator, no locks, no quorum. The cost is that the merge rule, not the application, decides what "concurrent" means, and the rule differs per type. ## How each Redis type behaves **Counters.** The naive replication of "the new value is 7" loses updates when two regions increment concurrently. Instead, increments replicate as **operations**: region A ships +1, region B ships +1, and both converge to +2 over the original value. This applies to INCR/INCRBY on strings, HINCRBY on hash fields, and sorted-set score increments. It is the strongest argument for active-active: distributed counting that actually adds up. **Collections (sets, hashes, sorted sets).** Elements are tracked with observed-remove semantics. A delete removes only the element instances the deleting replica had actually observed. So if region A adds `x` while region B deletes the `x` it had seen, the add survives — "add wins". This is deliberate: losing a newly created element is usually worse than keeping one that someone tried to remove. **Plain string values.** `SET k v` replaces an opaque blob. There is no meaningful merge of two different blobs, so the system resolves **last-write-wins** using the timing information it maintains across regions. One writer's value simply disappears. This is the single most important thing to state in an interview: active-active does *not* make concurrent overwrite of the same string safe — it makes it *deterministic*, which is not the same as *correct for your business*. **Expiry and deletes.** TTLs converge by keeping the longer expiry so a key does not vanish from under a region that just extended it. Deletes propagate as tombstones, which linger until garbage collection — part of why a CRDB uses noticeably more memory per key than a plain database. ## What the model does not give you - **No linearizability or global order.** Two regions can transiently disagree; there is no single ordering of operations that every observer sees. - **No read-your-writes across regions.** A write in Frankfurt is not visible in Virginia until replication delivers it. - **No cross-region atomicity.** MULTI/EXEC executes locally; WATCH-based optimistic concurrency only detects changes made in the local instance, so it cannot protect against a concurrent remote write. Scripts likewise run locally and must be written so their effects merge sensibly. - **No zero-RPO failover.** Replication is asynchronous, so an acknowledged local write can be lost if the region is destroyed before the write ships. Durability is per-region. - **No escape from application semantics.** "Deduct one from the balance" merges beautifully as a counter; "set the balance to 40" does not. Modelling matters more here than in any single-region deployment. ## Operational consequences **Cost.** Every region stores the entire dataset and pays cross-region bandwidth for every write, so a five-region CRDB is roughly five copies of the data plus a full write fan-out. **Metadata overhead.** CRDT bookkeeping and tombstones increase memory per key relative to an equivalent non-CRDB, and garbage collection of tombstones is a background cost. **Command restrictions.** Some commands behave differently or are restricted in a CRDB precisely because they cannot be merged meaningfully; check the supported-command matrix before assuming parity with a single-region database. **Session affinity helps a lot.** Pinning a user to one region for the duration of a session sidesteps most read-your-writes surprises without giving up the multi-region write path. ## Choosing it Active-active is right when you genuinely need local writes in multiple regions and your data model is dominated by mergeable operations — counters, tallies, set membership, presence, feature-flag exposure, session data with region affinity, shopping-cart style add/remove. It is the wrong tool when the truth of a value depends on reading it first and writing back a derived value across regions; that requires coordination, and no CRDT will supply it. The professional answer names both the mechanism and that boundary.
- Two regions concurrently run `SET user:1:name "A"` and `SET user:1:name "B"`. What is the outcome, and how would you avoid the situation?One value wins by last-write-wins and the other is silently discarded — all regions converge on the same survivor, but a write is lost. Avoid it by giving each writer its own key or hash field, by pinning a given entity's writes to one region through session or shard affinity, or by modelling the change as a mergeable operation rather than a whole-value overwrite.
- Can you use WATCH-based optimistic concurrency to guard a value in an active-active database?Not across regions. WATCH detects modifications visible to the local instance, so a concurrent write in another region does not abort your transaction — the two writes are simply merged afterwards by the type's CRDT rule. Within a single region it still works as usual. Any cross-region mutual exclusion needs a coordination mechanism outside the CRDB.
- Why do concurrent INCRs across regions produce the correct sum when concurrent SETs do not?Increments are replicated as operations, so each region ships a delta and applying both deltas in any order yields the same total — that is exactly the commutativity a CRDT needs. A SET replicates an opaque replacement value with no way to combine two different blobs, so the system can only pick a winner deterministically. The lesson is to express changes as deltas wherever the data type allows it.
Several branch offices each keep a tally sheet. If everyone records 'sold three more', the sheets can be merged by adding — order does not matter. If instead each branch writes down 'total is 40', merging is impossible and someone's number has to be thrown away. CRDTs make Redis record the first kind of statement wherever the type allows it.
saying these in an interview costs you the question
- Claiming active-active gives strong consistency or global ordering because conflicts are 'resolved'
- Assuming concurrent SETs on the same string are merged rather than one being discarded
- Believing WATCH/MULTI provides cross-region mutual exclusion in a CRDB
- Thinking asynchronous geo-replication implies zero data loss on regional failure
- Ignoring the extra memory from CRDT metadata and tombstones when sizing