Some Redis capabilities are absent from the open-source distribution because of how the system is constructed, not because of how it is packaged — for example accepting writes to the same dataset in more than one region simultaneously, serving cold values from flash instead of RAM, and moving shards between nodes without the connected client observing it. For each of those three, what would you actually have to build to provide it yourself on top of open-source Redis, and which workloads genuinely need none of them?
answer
- packaging gap vs construction gap
- merge rules live inside the data types
- concurrent SADD vs SREM → LWW loses intent
- swap stalls the single-threaded event loop
- invisible resharding = routing layer owns the connection
basics
~20 sThree gaps are structural, not packaging. Multi-master writes need merge rules built into every data type. RAM/flash tiering needs a storage engine under the shard with a promotion policy. Invisible resharding needs a routing layer owning the client connection. Single-region caching needs none of them.
solid answer
~50 sSeparate **packaging** gaps from **construction** gaps. Hosting, backups, dashboards, monitoring and support are packaging — buildable, or bought from any provider's managed Redis. Three are structural: - **Multi-master writes.** Convergence needs merge semantics inside each data type — what it means when a set element is added in one region and removed in another concurrently. That lives in the data structures, not in a replication process you can bolt on. (The active-active mechanics are covered by the dedicated question in this topic.) - **RAM/flash tiering.** Needs a storage engine beneath the shard with its own promotion and eviction policy. OS swap is not a substitute: one page fault stalls the single-threaded event loop. - **Invisible resharding.** Needs a routing layer that owns the client connection, so slots can move without the client seeing redirects or reconnects. The converse is honest: a single-region cache, a normal primary/replica or Cluster deployment, and automated failover via Sentinel or Cluster are fully served by open-source Redis.
code
text · 13 lines# Region A (t = 10:00:00.100) # Region B (t = 10:00:00.090)
SADD team:42 alice SREM team:42 alice
# An external replayer sees two opaque commands and picks the later
# timestamp -> alice present. Reverse the clock skew by 20ms and the same
# code yields alice absent. Neither outcome encodes intent.
#
# Correct convergence needs per-type rules decided INSIDE the data structure,
# e.g. add-wins with a per-element tombstone and a replica id:
# element alice: {added_by: A@ctr7, removed_by: B@ctr3}
# and the same for counters:
# INCRBY 1 in A and INCRBY 1 in B must sum to 2, which requires storing
# per-replica sub-counters rather than one integer.go deeper
Recall the distinction: hosting, backups and dashboards are packaging you can build or buy anywhere; multi-region writes, flash tiering and client-invisible resharding are different in kind. Naming the three and saying they need changes inside Redis itself is a good answer at this level.
Explain why the obvious shortcut fails for each: timestamp-based merging silently discards concurrent set or counter operations; OS swap stalls the single-threaded event loop; open-source Cluster resharding leaks MOVED/ASK redirects into every client. Also state that Sentinel and Cluster already cover ordinary failover.
Lead with where the missing machinery would have to live — inside the data types, beneath the shard, in front of the connection — and what building it would actually cost. Then check the requirement empirically: measure access skew before assuming tiering pays, and demand a written statement of why writes must be accepted in several regions.
Frame it as build-versus-adopt with attention to irreversibility. Adopting multi-master semantics changes the data model (counters decompose, sets acquire merge policies), so exiting means redesigning application logic, whereas adopting a proxy or tiering is comparatively reversible. Be equally clear that most workloads sit entirely inside what open-source Redis already covers.
## "Absent by construction" vs "absent by packaging" Most feature lists mix two very different kinds of gap. A **packaging gap** is something that exists as a wrapper around Redis: provisioning, backups, patching, dashboards, alerting, a support phone number, one-click replica creation. You can build every one of these against unmodified open-source Redis, and every cloud provider already sells a version of them. Nothing about Redis's internals stands in your way. A **construction gap** is something that requires changing the code *below* the level a client library, sidecar, or operator can reach. You cannot add it by configuration or by putting a process next to Redis; you would have to fork it. Three capabilities fall in this category, and knowing why each one is hard is the whole point of the question. ## Gap 1 — writes accepted in more than one place at once Ordinary Redis replication is single-master: one node accepts writes, replicas apply the stream. If two regions both accept writes to the same key, you need a rule for combining concurrent, independent changes — and the rule cannot be generic, because it depends on what the value *means*. Consider a set. Region A runs `SADD members alice`; concurrently region B runs `SREM members alice`. Neither happened "after" the other in any meaningful sense. A generic external merger sees two opaque commands and can only pick a winner by timestamp — last-writer-wins — which silently discards one user's intent. A correct answer needs per-type semantics: an add-wins or remove-wins rule for sets, a per-replica counter decomposition for integers so two concurrent `INCRBY 1` operations sum to 2 rather than collapsing to 1, and something defensible for strings, hashes, sorted sets and streams. Those rules require extra metadata stored *alongside every value* — vector clocks, per-replica sub-counters, tombstones with garbage-collection rules. That metadata has to be written by the same code that executes the command, which is why this is a change inside the data structures rather than a layer around them. Building it yourself means forking Redis and reimplementing conflict-free replicated types for every type you use, plus the tombstone GC, plus the replication protocol that carries the metadata. It is a multi-year project, not a config file. (The behaviour and caveats of ready-made active-active replication are covered by the dedicated question on it.) ## Gap 2 — a storage tier beneath RAM Redis holds values in the process heap. Keeping hot values in RAM and cold values on local flash means introducing a storage engine under the shard: an on-disk representation for values, an index from key to location, a **promotion policy** that decides when a flash-resident value must be pulled back into memory, and an eviction/demotion policy going the other way. The naive substitute — give the machine a big swap file and let the OS page — fails for a specific reason: Redis executes commands on a single thread, so a page fault taken while serving `GET` stalls every other client on that shard until the disk returns. The OS also pages at 4 KB granularity with no knowledge of key boundaries, so a hot key can sit on a cold page. A real tiering implementation performs the fetch off the command thread and keeps hot metadata (keys, expiries, and typically small values) resident so that the common path never touches disk. Building this yourself is a storage-engine project layered into a database that was designed on the assumption that memory access is free — and every latency guarantee Redis makes has to be re-derived once that assumption is gone. Its economics also depend entirely on access skew: without a small hot set, you get disk latency at RAM prices. ## Gap 3 — moving shards without the client noticing Open-source Redis Cluster *can* migrate slots while online, but it does so by making the client complicit: the client must understand the slot map, follow `MOVED` and `ASK` redirects, and refresh its topology. Every language's client implements that with its own bugs, and connections churn during a reshard. "Invisible" resharding means the client holds one stable endpoint and knows nothing about slots. That requires a routing layer that terminates the client connection and re-targets requests as topology changes — which means it now owns connection state, pipelining semantics, MULTI/EXEC affinity, blocking commands, and pub/sub subscriptions. Building it is a real distributed-systems component sitting on the hot path, and it introduces a hop that must be made highly available itself. (The proxy-versus-smart-client tradeoff is covered by the dedicated question on it.) ## The honest converse None of this is needed by most workloads. A single-region cache, a session or rate-limit store, a queue, a leaderboard — served entirely by open-source Redis. A standard primary/replica pair or a Cluster deployment, with automated failover through Sentinel or Cluster's own election, covers ordinary high availability; "we want automatic failover" is not a capability gap. Backups, TLS, ACLs, metrics and rolling upgrades are all present. If you cannot write down a requirement that specifically needs multi-region write acceptance, sub-RAM cost per gigabyte at verified skew, or client-invisible topology change, the structural gaps do not apply to you. ## How to answer this in an interview Name the three gaps, explain in one sentence each *where* the missing machinery would have to live (inside the data types, beneath the shard, in front of the connection), give the concrete failure of the obvious shortcut (last-writer-wins loses operations; swap stalls the event loop; redirects leak into every client), and then state plainly which workloads need none of it. The signal being tested is whether you can tell an architectural limitation from a pricing tier.
- Someone proposes achieving multi-region writes by running two independent Redis deployments and replaying each one's command stream into the other, resolving conflicts by timestamp. What breaks?Timestamp resolution is last-writer-wins, which is only correct for whole-value overwrites of independent keys. Concurrent structural operations — an element added in one region and removed in the other, or two independent increments of the same counter — lose one side's intent entirely, and the loss is silent. Correct convergence needs per-type merge rules plus per-value metadata written by the command execution path itself, which an external replayer cannot produce. Clock skew makes it worse, since 'later' is decided by unsynchronised wall clocks.
- A team's dataset is 2 TB with a strongly skewed access pattern and they want to cut RAM cost. Before reaching for tiering, what would you check?First measure the actual skew — what fraction of requests hit what fraction of the keyspace over a realistic window — because tiering's economics collapse without a small hot set. Then look at cheaper levers: value encoding and compression, shorter keys, hash-field packing for small objects, a lower TTL, or simply not storing cold data in Redis and falling back to the source of truth. Tiering is worth it when the hot set genuinely fits RAM and the cold tail is large and rarely touched.
- Is automated failover a capability gap in open-source Redis?No. Sentinel provides monitoring, leader election and promotion for primary/replica setups, and Redis Cluster performs replica promotion itself. What differs is the client experience — how long clients take to discover the new topology and whether the endpoint stays stable — which is a routing concern, not a missing failover mechanism. Treating 'we want automatic failover' as a structural gap is a category error.
Adding multi-master writes by putting a merge process next to Redis is like resolving two editors' conflicting changes to a document by keeping only the file with the later timestamp — you need to know what a sentence is to merge them, and only the editor holds that knowledge.
saying these in an interview costs you the question
- Calling active-active replication "bidirectional replication plus a conflict-resolution script" — it needs per-type merge metadata written inside the data structures.
- Claiming OS swap or a memory-mapped file gives you RAM/flash tiering, ignoring that a page fault blocks the single-threaded event loop for every client on the shard.
- Saying open-source Redis has no automated failover — Sentinel and Cluster both provide it.
- Assuming any managed Redis service implies multi-region write acceptance; managed hosting is a packaging concern and usually ships single-master replication.
- Claiming open-source Cluster resharding is already invisible to clients, when clients must follow MOVED/ASK redirects and refresh the slot map themselves.