skip to content

Redis Enterprise can keep part of a dataset on SSD rather than entirely in RAM (marketed as Auto Tiering, previously Redis on Flash). How does that work, what workload does it assume, and why is it not a durability feature?

level: seniorimportance: nice to knowfreq 18%

answer

  1. keys + metadata + hot values in RAM, cold values on local NVMe
  2. benefit depends entirely on access skew
  3. miss = flash read on a shard that executes one command at a time
  4. metadata never tiered → many tiny keys gain little
  5. tiering ≠ persistence; snapshots + replicas still required

basics

~20 s

Keys, indexes and hot values stay in RAM while cold values live on local NVMe SSD, fetched on access. It assumes a strongly skewed access pattern, since a miss costs an SSD read instead of a memory read. It is a cost-per-gigabyte optimisation, not persistence — durability still comes from snapshots and replication.

solid answer

~1 min

Auto Tiering splits a shard's data across two tiers on the same node: **RAM holds the keyspace metadata, key names and the hot values; local NVMe SSD holds cold values**, managed by an embedded storage engine. Access promotes a value back into RAM; cold values are demoted under memory pressure. The assumption is a **skewed working set** — a small fraction of keys serves most requests. When that holds, you can run a dataset several times larger than RAM at a much lower cost per gigabyte, with most requests still served from memory. When it does not hold, every access becomes a flash read: latency jumps from microseconds to hundreds of microseconds or more, and the throughput collapses because the shard is doing I/O in place of memory lookups. Critically, it is **not** persistence. The data on flash is a tier of the live dataset, not a durable copy; the SSD is local and treated as ephemeral. Durability still comes from snapshots/append-only persistence and replica shards, exactly as with an all-RAM database. Other constraints: key names and metadata always stay in RAM, so a dataset of billions of tiny keys is bounded by RAM regardless of tiering, and large values benefit more than small ones.

go deeper

for a junior

Know that hot values stay in RAM and cold values move to local SSD, and that it is about cost, not durability.

for a middle

Explain the RAM/flash ratio, that keys and metadata always stay in RAM, and why value size matters.

for a senior

Evaluate it against measured working-set skew and the latency SLO, and note the write-endurance and slower-restart consequences.

for a principal

Treat it as a cost-per-gigabyte lever that is only valid under a verified access-skew assumption, and compare it against sharding, eviction, or moving the cold tail to a different store entirely.

## Why the feature exists RAM is the dominant line item in a large Redis deployment. Many datasets, though, are strongly skewed: a small hot subset serves most of the traffic while a long cold tail is retained because it is occasionally needed. Paying DRAM prices for the cold tail is wasteful, and evicting it is not acceptable when it must remain queryable. Auto Tiering (formerly Redis on Flash) targets exactly that shape — it lowers cost per gigabyte by putting the cold tail on local NVMe SSD while keeping the hot path in memory. ## How the tiering works Within a shard, memory holds the structures that must be traversed on every request: the keyspace dictionary, key names, expiry information, and the values currently considered hot. Values that go cold are written down to an embedded storage engine on local flash and their in-memory representation is released. On access to a cold value, the shard reads it from flash and promotes it back to RAM; something else is demoted to make room. The policy is driven by access recency/frequency and by a configured RAM-to-flash ratio per database, so an operator effectively declares how much of the dataset should be resident. Two structural consequences follow: 1. **Metadata is never tiered.** Key names and the dictionary stay in RAM, so RAM consumption scales with the *number* of keys no matter how cold they are. A billion 20-byte values gains little from tiering; a hundred million multi-kilobyte values gains a lot. 2. **Value size drives the benefit.** Bigger values mean more bytes moved off DRAM per unit of retained metadata, and better amortisation of each flash read. ## The workload assumption, stated plainly The entire economic argument rests on **skew**. If 90% of requests hit values that are resident in RAM, the average latency stays close to an all-RAM database and only the remaining 10% pay a flash read. If access is uniform over a dataset several times larger than RAM, nearly every request becomes an SSD read: latency moves from single-digit microseconds to hundreds of microseconds or milliseconds, tail latency degrades much further, and the shard — which executes commands one at a time — is now blocked on I/O where it used to be blocked on nothing. That means the honest evaluation criteria are: measure the working-set size and the hit ratio you would get at the intended RAM ratio, and check whether your latency SLO tolerates the miss path at your miss rate. Write-heavy workloads that constantly dirty cold data are the least suitable, because they generate flash writes as well as reads and consume device endurance. ## Why it is not durability This is the most common misconception and the one an interviewer is probing. The flash tier holds part of the **live dataset**, not a recovery copy. It is local to the node, it is not designed to be replayed after a crash, and losing the node loses that shard's data exactly as it would in an all-RAM configuration. Durability and availability come from the same mechanisms as always: snapshot and/or append-only persistence configured per database, and replica shards on other nodes for failover. Tiering is orthogonal — you can and normally do enable both, and one does not substitute for the other. ## Operational notes - **Local NVMe is required.** Network-attached storage defeats the latency budget the design assumes. - **Provisioning is per database**, expressed as a RAM/flash ratio, so different databases on the same cluster can make different choices. - **Restart and failover are slower** than for an all-RAM database of the same *hot* size, because the shard has more total data to bring back into a serving state. - **Monitor the resident-hit ratio, not just memory usage.** Memory looking comfortable while hit ratio drops is the early signal that the working set has outgrown the RAM tier. ## When to choose it Good candidates: large caches or key-value stores with pronounced hot/cold skew and reasonably sized values — user profiles, product catalogues, feature stores, session archives — where the cold tail must remain accessible and the SLO can absorb an occasional flash read. Poor candidates: latency-critical paths with tight tails, uniformly accessed datasets, huge counts of tiny keys, and write-dominated workloads over cold data. As with the rest of the Enterprise capability set, the interesting answer is not that the feature exists but that its benefit is entirely contingent on an access-pattern assumption you should verify before adopting it.

  • Your dataset is 2 TB with nearly uniform access and a 5 ms p99 requirement. Is Auto Tiering appropriate?
    No. Uniform access means the RAM tier gives you no hit-rate advantage, so most requests become flash reads and both average and tail latency degrade sharply. Tiering only pays off when a small hot subset absorbs most traffic. With uniform access you either provision enough RAM, shard across more nodes, or reconsider whether Redis is the right store for that dataset.
  • Does enabling Auto Tiering mean you can turn off persistence, since the data is already on SSD?
    No. The flash tier is part of the live dataset, held in a local storage engine that is not intended as a crash-recovery copy, and it disappears with the node. Durability still comes from snapshot/append-only persistence and from replica shards on other nodes. Tiering is a cost-per-gigabyte optimisation that is completely orthogonal to durability.

A reference library keeps the frequently borrowed titles on the open shelves and the rest in the basement stacks. It works beautifully while most requests are for shelf titles; if every reader wants something from the basement, the librarian becomes the bottleneck no matter how large the building is.

saying these in an interview costs you the question

  • Treating the flash tier as persistence or as a replacement for replicas
  • Assuming latency is unchanged because 'SSDs are fast'
  • Expecting big savings on a dataset made of billions of tiny keys, ignoring that metadata stays in RAM
  • Enabling it for a uniformly accessed dataset and expecting the hit ratio to save you
  • Believing it can use network-attached storage without consequences

context