When does putting Amazon DynamoDB Accelerator (DAX) in front of a DynamoDB table actually help, and which reads does it not accelerate?
answer
- it speaks the same API
- two caches, two different keys
- strong reads walk right past it
- out-of-band writes go unnoticed
- fixed nodes versus saved reads
basics
~20 sDAX is a write-through, in-VPC cache cluster for DynamoDB that turns repeated reads of the same items into microsecond responses. It helps read-heavy workloads with strong key reuse, and does nothing for strongly consistent reads, writes, or read patterns with poor locality.
solid answer
~50 sDAX is a managed cache cluster that speaks the DynamoDB API, so an application adopts it by swapping in the DAX client rather than rewriting queries. It keeps two caches: an *item cache* for `GetItem` and `BatchGetItem` results, and a *query cache* for `Query` and `Scan` result sets, each with a short configurable TTL. Writes go through DAX to the table and update the item cache, so DAX-mediated writes stay coherent. It earns its cost when the workload is read-dominated, the same items are read repeatedly, and read capacity or single-digit-millisecond latency is the constraint. It does nothing useful when reads have poor locality, when the workload is write-heavy, or when reads must be strongly consistent — DAX passes those straight through to DynamoDB. It also lives inside a VPC, so callers must have network access to the cluster.
go deeper
Know that DAX is a managed cache that sits in front of a DynamoDB table, speaks the same API, and mainly speeds up repeated reads of the same items.
Explain the item cache versus the query cache and their different keys, why writes are write-through, and that strongly consistent reads bypass the cache entirely.
Judge whether it earns its keep: estimate the hit rate from the real read pattern, weigh fixed node cost against saved read units, and call out out-of-band writes as the coherence risk.
Decide where caching belongs in the architecture at all — item-level at the database, response-level at the edge, or not needed after a data-access change — and set the staleness contract the whole system is designed around.
## What DAX is DynamoDB Accelerator is a fully managed, in-VPC cache cluster dedicated to a DynamoDB workload. It sits between the application and the table, exposes the DynamoDB data-plane API, and returns cached results in microseconds where DynamoDB itself answers in single-digit milliseconds. A cluster is made of nodes across Availability Zones, with a primary that handles writes and read replicas; it is provisioned with a node type and a replication factor, and priced per node-hour rather than per request. Adoption is deliberately cheap in code terms: you replace the standard DynamoDB client with the DAX client from the same SDK family and point it at the cluster endpoint. Calls that DAX cannot serve are forwarded to DynamoDB transparently. ## The two caches DAX maintains two logically separate caches with different keys. The **item cache** holds the results of `GetItem` and `BatchGetItem`, keyed by the item's primary key. A hit returns the item without touching DynamoDB and without consuming table read capacity. The **query cache** holds whole result sets of `Query` and `Scan`, keyed by the request parameters. This is the subtler one: because the key is the request, a query cached under one set of parameters is not invalidated when an underlying item changes. A write through DAX updates the item cache entry for that item, but the query cache is only refreshed by its TTL. Stale query results are therefore a normal, expected behaviour, not a bug — the TTL (five minutes by default, configurable) is the knob that bounds the staleness. ## Where it does nothing **Strongly consistent reads.** A request with `ConsistentRead=true` is passed straight to DynamoDB and its result is not cached. Any code path that must read its own writes is unaffected by DAX and still pays full read capacity. **Writes.** DAX is write-through, meaning a write goes to DynamoDB first and the item cache is updated on success. That adds a small amount of latency to writes relative to talking to the table directly. A write-heavy table gains nothing and pays for the cluster. **Cold or scattered reads.** A cache only helps when the same data is read repeatedly. A workload that reads each item once — an export, a fan-out over unique keys, a long tail with no popular items — will see a low hit rate and a bill for nodes that mostly forward requests. **Writes that bypass DAX.** This is the failure mode teams actually hit. If some services write through the DAX client while others (or a Lambda triggered by a stream, or a batch job, or a console edit) write directly to the table, DAX has no way to know the item changed and will serve the stale value until its TTL expires. Coherence is only as good as the discipline that all writes go through the cluster. ## The network and operational shape DAX clusters are VPC resources reached over a cluster endpoint on their own port, secured with a security group. That has real consequences: a Lambda function that is not attached to the VPC cannot reach DAX at all, and neither can anything outside the VPC. The cluster also assumes an IAM role that grants it access to the table, which is a second place authorisation can be misconfigured — the caller needs permission to talk to DAX, and DAX needs permission on the table. Because it is provisioned by node, DAX is a fixed cost that must be justified against the variable cost it removes. The arithmetic is straightforward: estimate the hit rate from your read pattern, multiply the saved reads by the read-unit price, and compare with the node-hours. For a genuinely hot, read-heavy table the saving is large; for a lukewarm one it is negative. ## The alternatives worth naming DAX is not the only cache in front of DynamoDB. A general-purpose cache lets you cache computed responses rather than raw items, which is often the higher-leverage layer, at the cost of writing the cache logic yourself. Reading through an eventually consistent read, or restructuring so that the hot data is read once per request rather than many times, sometimes removes the need for a cache entirely. DAX's differentiator is that it requires almost no application change and stays coherent for writes that flow through it. ## How to answer Describe DAX as a write-through, API-compatible cache with an item cache and a query cache, then immediately state the three places it does not help — strongly consistent reads, writes, and poor key locality — and the coherence caveat about out-of-band writes. Finishing with the cost comparison (fixed node-hours against saved read units) turns a feature description into an engineering judgment.
- An application uses DAX but one background job writes directly to the DynamoDB table. What goes wrong?DAX never learns about those writes. Its item cache keeps serving the previous value until the entry's TTL expires, so readers see stale data for up to that window with no error anywhere. Either route every writer through the DAX client, or shorten the TTL to a staleness you can tolerate and document it. Consistency here is a convention, not an enforced guarantee.
- Why can a DAX query cache return a result that contradicts the item cache?They are keyed differently. The item cache is keyed by primary key and is updated by write-through, so it tracks individual items. The query cache is keyed by the Query or Scan request parameters and is only refreshed when its TTL expires, so a result set can still contain a pre-write view of an item that the item cache has already updated.
- When would you choose a general-purpose cache over DAX for a DynamoDB workload?When what you want to cache is not a raw item. A general-purpose cache can hold an assembled API response, an aggregate, or a fragment shared across many items, cutting far more work than an item lookup. It also serves callers outside the VPC and is not tied to one database. The tradeoff is that invalidation becomes your problem, where DAX handles it for writes that pass through it.
saying these in an interview costs you the question
- DAX accelerates strongly consistent reads too
- DAX invalidates the query cache whenever an item changes
- DAX makes writes faster as well as reads
- Any client can reach a DAX cluster over the public internet
- A cache always pays for itself on a read-heavy table