skip to content

ElastiCache

Managed Redis/Valkey or Memcached with replication groups, automatic failover, and a serverless option that hides node sizing entirely. Interviewers ask which engine I would pick and what happens to the cache — and to the database behind it — when a node dies.

part ofAWSoverview, primer and where to startread it →
on this pageshow

questions

6

Amazon ElastiCache lets you run Valkey, Redis OSS, or Memcached as the cache engine. What does each option give you operationally, and for which workload would you actually choose Memcached?

level: juniorimportance: must knowfreq 72%

answer

  1. one engine replicates, one does not
  2. multi-threaded versus feature-rich
  3. failover, snapshots, auth on one side
  4. losing a Memcached node loses its keys
  5. Valkey is a fork, not AWS-proprietary

basics

~10 s

Valkey and Redis OSS on ElastiCache add replicas, automatic failover, snapshots and authentication; Memcached has none of those but is multi-threaded and simple. Choose Memcached only for a plain, disposable, horizontally sharded key-value cache.

solid answer

~50 s

ElastiCache offers three engines. Valkey and Redis OSS run as a **replication group**: a primary with read replicas, Multi-AZ automatic failover, backup and restore, and support for TLS, AUTH tokens and RBAC. Memcached runs as a flat set of independent nodes — no replication, no failover, no snapshots — where the client shards keys across nodes and ElastiCache's Auto Discovery keeps the node list current. Memcached executes commands on multiple threads, so one large node uses all its cores, while Valkey and Redis OSS execute commands on a single thread per shard. In practice the default choice is Valkey or Redis OSS, and the reason is the operational feature set, not raw speed. Memcached is the right pick only when the cache is genuinely disposable, values are opaque blobs, and losing a node's share of the data costs nothing but a few extra database reads.

code

bash · 17 lines
bash
# Memcached: a flat set of independent nodes, no replicas, no failover
aws elasticache create-cache-cluster \
  --cache-cluster-id sessions-mc \
  --engine memcached \
  --cache-node-type cache.m7g.large \
  --num-cache-nodes 3

# Valkey: a replication group with replicas and automatic failover
aws elasticache create-replication-group \
  --replication-group-id app-cache \
  --replication-group-description "application cache" \
  --engine valkey \
  --cache-node-type cache.m7g.large \
  --num-node-groups 1 \
  --replicas-per-node-group 2 \
  --automatic-failover-enabled \
  --multi-az-enabled

go deeper

for a junior

Know the headline split: Valkey and Redis OSS support replicas, failover and backups; Memcached does not. Be able to say plainly that a Memcached node failure loses that node's cached data.

for a middle

Explain how each engine scales — Memcached shards in the client via Auto Discovery, Valkey and Redis OSS shard server-side with cluster mode and add replicas for reads — and why one is multi-threaded and the other is not.

for a senior

Justify the choice from the blast radius of a node loss: whether the source of truth can absorb that share of traffic, and whether the cache needs backups or authentication to be operable.

for a principal

Own the standard for the org. Decide whether teams get one blessed engine, what the migration path from Redis OSS to Valkey looks like across a fleet, and how the pricing difference nets against the retest cost.

## What ElastiCache is, and what the engine choice actually decides ElastiCache is a managed in-memory cache that runs inside your VPC. AWS provisions the nodes, patches the engine, monitors them, replaces failed ones, and — for some engines — fails over automatically. What you still choose is the **engine**, and that choice decides which of those managed behaviours are even available to you. Today ElastiCache offers Valkey, Redis OSS and Memcached. The interview version of this question is rarely about which engine is faster. It is about whether you know that one engine has an availability story and the other does not. ## Memcached on ElastiCache A Memcached cluster is a set of independent nodes that know nothing about each other. Concretely: - **No replication and no failover.** There are no replicas, so there is nothing to promote. When a node fails, ElastiCache replaces it with an empty one and that node's share of the cache is gone. - **No backup and restore.** There is no snapshot of a Memcached cluster to export or restore from. - **Values are opaque.** Memcached stores byte strings keyed by a name. There are no server-side collections, no pub/sub, no server-side scripting. - **Multi-threaded.** A Memcached node serves commands on several threads, so a large multi-core node type is used efficiently by a single node. - **Sharding lives in the client.** ElastiCache publishes a *configuration endpoint*, and a client that supports **Auto Discovery** uses it to learn the current node list; the client then hashes each key to a node itself. Adding or removing a node redistributes some keys and those entries are simply missed until they are re-populated. ## Valkey and Redis OSS on ElastiCache These engines are deployed as a **replication group**: one primary per shard plus read replicas. That single structural difference is what unlocks the managed features: - **Multi-AZ with automatic failover** — put a replica in another Availability Zone and ElastiCache promotes it when the primary fails, keeping the cached data. - **Backup and restore** — scheduled or manual snapshots, restorable into a new cluster and exportable to S3. - **Read scaling** — replicas serve reads through a reader endpoint. - **Horizontal write scaling** — cluster mode splits the keyspace across shards. - **Security controls** — encryption in transit, an AUTH token, and user-level RBAC. - **A rich command set** — server-side collections, pub/sub, streams and scripting, so the cache can hold more than opaque blobs. The cost is that command execution on one shard is effectively single-threaded, so a single node's throughput ceiling is one core's worth of work; you add capacity by adding shards or replicas rather than by buying more cores in one box. ## Valkey versus Redis OSS Valkey is the community fork of Redis OSS created under the Linux Foundation after Redis changed its licence. On ElastiCache it is protocol-compatible — existing clients and commands work — and AWS prices it below Redis OSS for equivalent capacity. For a new cluster, Valkey is the usual default unless a specific client, module or engine version pins you to Redis OSS. It is not an AWS-proprietary engine, which is a common misconception worth correcting out loud. ## When Memcached is genuinely the right answer Say Memcached when all of these hold: 1. The cached data is regenerable from the source of truth and losing a slice of it is a cost blip, not an incident. 2. Values are opaque strings you serialize yourself, with no need for server-side structures. 3. You want maximum throughput per node on a memory-bound, CPU-light workload and want the node's cores used. 4. Nothing needs failover, backups, or per-user authentication of the cache tier. A session store, a rate-limiter, a leaderboard, a queue, or any cache whose loss would stampede the database into an outage all point the other way. ## The failure mode candidates miss The most common wrong answer is "Memcached is faster, so use it for caching." The throughput difference is real but rarely the constraint; the availability difference is what shows up at 3 a.m. If a node loss means the database suddenly absorbs that node's whole read share, you wanted an engine that can promote a replica with the data still in it.

  • Memcached nodes are added and removed without any cluster protocol between them. What does ElastiCache give you to make that manageable from the client side?
    A configuration endpoint plus Auto Discovery. A client that supports it queries that endpoint for the current node list and refreshes it, instead of having the node hostnames baked into config. The client still hashes keys across nodes itself, so adding or removing a node redistributes part of the keyspace and those entries miss until they are repopulated.
  • Why does AWS list Valkey alongside Redis OSS, and how would you choose between them for a new cluster?
    Valkey is the Linux Foundation fork created after Redis changed its licence. On ElastiCache it is protocol-compatible, so existing clients work unchanged, and AWS prices it below Redis OSS. For a new cluster Valkey is the default unless a client library, an engine version, or a feature you depend on pins you to Redis OSS.
  • If Memcached loses a node's data on failure, why do people still run it in production?
    Because for a purely regenerable cache the loss is a bounded cost — a burst of extra reads against the source of truth — not a correctness problem. In exchange you get a multi-threaded engine that uses a big node's cores well and a very simple operational model with no failover semantics to reason about.

saying these in an interview costs you the question

  • Says Memcached is faster so it is always the better cache
  • Thinks ElastiCache for Memcached supports read replicas and failover
  • Believes a Memcached cluster can be snapshotted and restored
  • Calls Valkey an AWS-proprietary engine rather than a community fork
  • Picks Redis OSS purely out of habit without naming a feature it needs

context

open as a page

In ElastiCache for Valkey or Redis OSS, what is the difference between a replication group with cluster mode disabled and one with cluster mode enabled, and what does that choice force on the application's client library?

level: middleimportance: must knowfreq 58%

basics

~20 s

Cluster mode disabled means one shard holding the whole keyspace, scaled up by node size. Cluster mode enabled splits the keyspace across many shards, each with its own primary, and requires a cluster-aware client that follows slot redirections.

open as a page

An ElastiCache for Valkey replication group with cluster mode disabled exposes a primary endpoint and a reader endpoint. What does each one resolve to, and what happens if the application sends its writes to the reader endpoint?

level: middleimportance: should knowfreq 50%

basics

~20 s

The primary endpoint is a DNS name that always tracks the current primary, including after a failover. The reader endpoint resolves across the read replicas. Writes sent to a reader are rejected by the engine with a read-only error.

open as a page

An ElastiCache for Valkey replication group is Multi-AZ with automatic failover enabled, and its primary node fails. Walk through what ElastiCache does, what the application sees, and what you must have configured beforehand for the recovery to be clean.

level: seniorimportance: should knowfreq 46%

basics

~20 s

ElastiCache detects the failure, promotes a replica, and repoints the primary endpoint's DNS at it. Clients see dropped connections and errors for tens of seconds, and any write acknowledged but not yet replicated is lost, because replication is asynchronous.

open as a page

You are choosing between ElastiCache Serverless and a node-based ElastiCache cluster you size yourself for a new service. How do you make that call, and what do you give up either way?

level: principalimportance: should knowfreq 34%

basics

~20 s

Decide on workload shape and control. Serverless removes node sizing, shard layout and capacity planning, and bills for data stored plus processing consumed — good for spiky or unknown traffic. Node-based clusters cost less at steady high load and keep node-level tuning.

open as a page

An ElastiCache for Valkey cluster holds session data for a web application. What mechanisms does ElastiCache give you to control who can connect and to protect the traffic, and which layer does an IAM policy actually govern?

level: middleimportance: nice to knowfreq 38%

basics

~20 s

Four layers: the cluster is VPC-only behind security groups, encryption in transit protects the wire, an AUTH token or RBAC users authenticate clients, and RBAC access strings limit commands and key patterns. IAM policies govern the management API, not the data commands.

open as a page