skip to content

An ElastiCache for Valkey cluster holds session data for a web application. What mechanisms does ElastiCache give you to control who can connect and to protect the traffic, and which layer does an IAM policy actually govern?

level: middleimportance: nice to knowfreq 38%

answer

  1. four layers, not one control
  2. private subnet plus security group first
  3. the token needs TLS to exist
  4. per-user access strings limit commands
  5. IAM governs the API, not the commands

basics

~20 s

Four layers: the cluster is VPC-only behind security groups, encryption in transit protects the wire, an AUTH token or RBAC users authenticate clients, and RBAC access strings limit commands and key patterns. IAM policies govern the management API, not the data commands.

solid answer

~60 s

ElastiCache stacks four controls. **Network**: the cluster has no public endpoint, lives in a subnet group in private subnets, and its security group should admit only the application's security group on the engine port. **Transport**: enable encryption in transit so clients connect over TLS, and connect using the DNS endpoint so certificate validation works. **Authentication**: either an AUTH token — which requires in-transit encryption — or, on engine version 6 and later, **RBAC**, where you create ElastiCache users and user groups and attach them to the replication group; recent engine versions also support IAM authentication, where the client presents a short-lived IAM-derived token. **Authorization**: an RBAC user's access string limits which commands it may run and which key patterns it may touch, so a read-only consumer cannot flush the cache. The distinction that catches people is that an IAM policy on `elasticache:*` governs the *management* API — creating, modifying and deleting clusters — not what a connected client is allowed to execute, unless you are specifically using IAM authentication.

code

bash · 6 lines
bash
# Rotate the AUTH token without recreating the cluster
aws elasticache modify-replication-group \
  --replication-group-id app-cache \
  --auth-token "$NEW_TOKEN" \
  --auth-token-update-strategy ROTATE \
  --apply-immediately

go deeper

for a junior

Know that the cluster lives in a VPC behind a security group, that traffic can be encrypted with TLS, and that clients can be made to authenticate rather than connecting anonymously.

for a middle

Explain each layer and how they interact — why the AUTH token requires in-transit encryption, how RBAC users and user groups attach to a replication group, and what an access string restricts.

for a senior

Draw the control-plane versus data-plane line clearly, plan a rotation and a TLS enablement that do not break running clients, and justify per-service identities for a shared cache holding credential material.

for a principal

Set the org-wide baseline: whether any cache may run without authentication, whether IAM authentication is mandated for new clusters, and how cache credentials are issued, stored and rotated across teams.

## Four layers, not one ElastiCache security questions go wrong when a candidate names only one control. There are four, and they answer different questions. ### 1. Network reachability An ElastiCache cluster is created inside a VPC using a **cache subnet group** and has no public endpoint. Reachability is governed by the security group attached to the cluster: the correct rule allows inbound traffic on the engine port **from the application's security group**, not from a CIDR block covering the whole VPC. This is the cheapest and strongest control, and it is where most real-world protection comes from. It is also not sufficient on its own. "It's private, so it doesn't need authentication" is a genuinely dangerous position: anything that lands in the VPC — a compromised instance, a misconfigured task, a developer's bastion — reaches the cache with full command access. ### 2. Encryption in transit Enabling **encryption in transit** makes clients speak TLS to the cluster. Two practical consequences: - Clients must be configured for TLS and must connect by the ElastiCache **DNS endpoint**, because certificate validation is against that hostname; connecting by IP breaks validation. - It is a prerequisite for the AUTH token — a shared secret sent over a plaintext connection would be pointless, so ElastiCache refuses that combination. ElastiCache Serverless has encryption in transit on by default, which removes the choice entirely. ### 3. Authentication Three options, in rough order of maturity: - **AUTH token** — a shared secret the client sends on connect. Simple, but it is one credential for everyone, so rotation is a fleet-wide event. ElastiCache supports rotating it without recreating the cluster: `modify-replication-group` accepts `--auth-token` together with `--auth-token-update-strategy ROTATE`, which allows the old and new token to be accepted during the migration window. - **RBAC users** (engine 6 and later) — you create ElastiCache **users**, group them into a **user group**, and attach the group to the replication group. Each user has its own credential, so an application, an analytics job and an on-call human can carry different identities. - **IAM authentication** (recent engine versions) — a user created with an IAM authentication mode, where the client presents a short-lived token derived from its IAM identity instead of a stored password. This removes a long-lived secret from the application entirely, which is the direction to prefer for new work. ### 4. Authorization RBAC is not only authentication. Each user carries an **access string** that declares which commands and which key patterns it may use. That is what lets you hand a reporting job an identity that can read `report:*` and nothing else, and keep destructive administrative commands away from application credentials. Without RBAC, every authenticated client can do everything. ```bash aws elasticache create-user \ --user-id app-reader --user-name app-reader \ --engine valkey --access-string "on ~report:* +@read" \ --passwords "$READER_SECRET" aws elasticache create-user-group \ --user-group-id app-users --engine valkey \ --user-ids default app-reader aws elasticache modify-replication-group \ --replication-group-id app-cache \ --user-group-ids-to-add app-users --apply-immediately ``` Note the `default` user in the group — leaving it able to authenticate without a credential defeats the whole exercise, so it should be locked down as part of adopting RBAC. ## Where IAM fits, and where it does not This is the discriminating part of the answer. An IAM identity policy allowing `elasticache:CreateReplicationGroup`, `elasticache:ModifyReplicationGroup` or `elasticache:DeleteReplicationGroup` controls the **control plane** — who may create, reconfigure or destroy clusters. It says nothing about what a client that has already opened a TCP connection may execute. Data-plane authorization comes from RBAC access strings; the only bridge between the two worlds is IAM authentication, where an IAM principal is what a data-plane user authenticates *as*. Encryption **at rest** is a separate axis again, configured with a KMS key at creation and covering snapshots and on-disk data — worth naming so the interviewer knows you have not confused it with in-transit encryption, but it protects a different threat. ## What a good answer covers for session data specifically Session tokens in a cache are credential material. That argues for: private subnets and a tight security group; TLS on; an authenticated identity per application rather than one shared token; RBAC access strings that keep each service inside its own key prefix; and rotation that does not require a maintenance window. It also argues for treating the cache as one more place secrets live, which means the credential itself belongs in Secrets Manager rather than in an environment variable baked into an image.

  • An engineer argues the cache needs no authentication because it sits in a private subnet. How do you answer?
    Network isolation is a strong control but a single one. Anything that gets a foothold in the VPC — a compromised instance, a mis-scoped task role, a bastion — reaches the cache with unrestricted command access, including destructive ones. Authentication plus per-user access strings turns that from total compromise into a bounded one, and costs almost nothing to enable.
  • What breaks when you turn on encryption in transit for an existing cluster, and how do you avoid it?
    Clients that are not configured for TLS stop connecting, and clients that connect by IP address rather than the ElastiCache DNS endpoint fail certificate validation. Roll it out by shipping TLS-capable, endpoint-addressed client configuration first, verifying in a non-production group, and only then changing the cluster — the client change must lead.
  • Why is RBAC preferable to a single AUTH token for a cache shared by several services?
    One shared token gives every service identical, unrestricted access and makes rotation a fleet-wide coordination problem. RBAC gives each service its own user with an access string scoped to the commands and key patterns it actually needs, so one leaked credential is contained and can be rotated on its own.

saying these in an interview costs you the question

  • Says an IAM policy controls which commands a client may run
  • Claims a private subnet makes authentication unnecessary
  • Thinks an AUTH token works without encryption in transit
  • Confuses encryption at rest with encryption in transit
  • Adopts RBAC but leaves the default user unrestricted

context