skip to content

A team wants active-active multi-Region writes for a DynamoDB-backed service and proposes global tables. What do global tables actually give them, what do they not, and how should that shape the design?

level: principalimportance: should knowfreq 36%

answer

  1. every replica takes writes
  2. the clock decides who wins
  3. one Region owns each entity
  4. replication copies your mistakes too
  5. the bill multiplies by replica count

basics

~20 s

Global tables replicate a DynamoDB table across Regions with writes accepted in every replica, converging asynchronously with last-writer-wins conflict resolution. They do not give cross-Region transactions, cross-Region read-after-write, or a backup — the application must be designed to tolerate convergence.

solid answer

~50 s

Global tables turn one table into a set of Regional replicas that all accept writes, propagating changes to each other over DynamoDB Streams. Replication is typically sub-second but is asynchronous and eventual, and simultaneous writes to the same item in two Regions are resolved by last writer wins on the observed timestamp — one of the two updates is silently discarded. Within a Region reads and conditional writes behave normally; across Regions there is no strong consistency by default and no distributed transaction. So the architecture has to route a given entity's writes to one Region where possible, model updates so that losing a concurrent one is acceptable (append-only events, per-Region attributes, CRDT-like counters rather than read-modify-write), and treat replicas as availability, not durability — a bad write replicates everywhere in a second, which is why PITR still matters. Cost roughly multiplies with the number of replicas, since replicated writes are billed in each Region.

code

bash · 6 lines
bash
aws dynamodb update-table \
  --table-name Orders \
  --replica-updates '[{"Create": {"RegionName": "eu-west-1"}}]'

aws dynamodb describe-table --table-name Orders \
  --query 'Table.Replicas[].{Region:RegionName,Status:ReplicaStatus}'

go deeper

for a junior

Know that a global table replicates a DynamoDB table into other Regions and that every replica accepts both reads and writes.

for a middle

Explain that replication is asynchronous over the table's stream, that conflicts resolve by last writer wins, and that strongly consistent reads apply only within a single Region.

for a senior

Show how you would operate it: monitor replication latency before failing over, keep PITR per replica, and identify the read-modify-write and uniqueness patterns that break across Regions.

for a principal

Own the decision itself — whether the requirement is latency, availability or disaster recovery, whether pinned writes with read replicas would meet it more cheaply, and what data-model rules the organisation adopts so convergence is safe by construction.

## What the feature is A DynamoDB **global table** is a single logical table with replicas in several Regions. Every replica is writable, every replica holds the full data set, and DynamoDB propagates each change to the other replicas asynchronously using the table's own stream mechanism. Adding a Region is an `UpdateTable` call: ```bash aws dynamodb update-table \ --table-name Orders \ --replica-updates '[{"Create": {"RegionName": "eu-west-1"}}]' ``` Applications talk to their local Regional endpoint with the ordinary API. Nothing in the data plane changes. ## What you genuinely get **Local read and write latency.** Users in Europe write to the Frankfurt replica rather than crossing an ocean. For a globally distributed user base this is the main reason the feature exists. **Regional failure isolation.** If a Region becomes unavailable, traffic can be shifted to another replica that already has the data and already accepts writes. Recovery time is a DNS or routing decision, not a restore. **Operational simplicity.** Replication is managed: no pipeline to run, no consumer to monitor, and schema-free tables mean no migration coordination between Regions. ## What you do not get, and must design around **Cross-Region consistency.** Replication is eventual. Typical propagation is well under a second, but nothing bounds it, and there is no cross-Region read-after-write: a client that writes in one Region and immediately reads in another may not see its own write. Strongly consistent reads work only against the local replica. (AWS has since added an optional multi-Region strong consistency mode with tighter Region-configuration constraints; the classic, default behaviour and the one interviews assume is eventual.) **Conflict handling beyond last-writer-wins.** If the same item is updated in two Regions close enough in time, DynamoDB keeps the write with the later timestamp and discards the other — silently, with no error and no merge hook. Read-modify-write patterns are therefore unsafe across Regions: two increments of a counter can collapse into one. **Cross-Region transactions or conditions.** `TransactWriteItems` and condition expressions are evaluated within a single Region. A conditional write that succeeds in one Region cannot know about a competing write in another that has not arrived yet. Anything depending on global uniqueness — claiming a username, reserving the last unit of stock — cannot be enforced by the table alone. **A backup.** Replication faithfully copies mistakes. A bad deployment that deletes items removes them everywhere within a second. Point-in-time recovery and on-demand backups remain necessary and are configured per replica. ## The design patterns that follow **Pin an entity to a home Region.** The most robust pattern is to make active-active a property of the *fleet* rather than of each item: route all writes for a given customer, tenant or account to one Region — via latency-based routing plus a stable mapping, or an explicit home-Region attribute — and let the other replicas serve reads and stand ready for failover. Conflicts then cannot occur in normal operation, and multi-Region write capability is exercised only during a failover, when a brief conflict window is an acceptable price. **Model updates so a lost write is tolerable.** Append immutable events keyed by Region and identifier rather than mutating a shared record. Keep per-Region attributes so two Regions never write the same attribute. Reconstruct aggregates by reading the events instead of maintaining a counter that both Regions increment. **Keep uniqueness out of the table.** Enforce globally unique claims with a single authoritative Region for that operation, or design them away with client-generated identifiers. ## Cost and operations Each replica stores the full data set and is billed for its own storage. Writes are charged as replicated write units in every Region, so a three-Region table costs roughly three times the write bill of a single-Region one, plus cross-Region data transfer. Streams must be enabled with both images. Adding or removing a replica is an online operation but takes time proportional to the data size, so a Region cannot be added mid-incident as an escape hatch. Monitoring should include the replication latency metric so you know how far behind a replica is running before you fail over to it. ## The honest recommendation Before agreeing to global tables, ask what problem is being solved. If it is disaster recovery, a single Region with PITR and a documented restore may meet the objective at a fraction of the cost and complexity. If it is read latency for a global audience, replicas with pinned writes give nearly all the benefit with none of the conflict risk. Truly concurrent multi-Region writes to the same item are the case where global tables shine and also the case that demands the most from the data model — and that is the tradeoff a lead is expected to name out loud rather than discover later.

  • If replicas converge in under a second, why is last-writer-wins still a real problem?
    Because the window is short, not zero, and the failures it produces are silent. Two Regions updating the same item within that window lose one update with no error, so the bug appears as quiet data drift rather than a visible failure. It also breaks read-modify-write outright: concurrent increments in two Regions can collapse into a single increment, and nothing in the table records that it happened.
  • A team argues global tables remove the need for point-in-time recovery. How do you respond?
    Replication is not backup. Every delete, every bad migration, every buggy batch job propagates to all replicas within about a second, so replicas defend against Regional loss and against nothing else. PITR is what protects against the far more common failure — your own code. Keep PITR enabled per replica and know the restore path, which creates a new table rather than restoring in place.
  • How would you route writes so that conflicts cannot occur in normal operation?
    Give each entity a home Region and route all of its writes there — for example a home-Region attribute resolved at the edge, or a deterministic mapping from tenant to Region — while every replica serves local reads. Concurrent writes to the same item then only happen during a failover, which is a bounded and understood window rather than a permanent property of the design.

saying these in an interview costs you the question

  • Global tables give strongly consistent reads across Regions
  • Conflicts are surfaced to the application to resolve
  • A global table replaces backups and point-in-time recovery
  • Transactions and condition expressions work across Regions
  • Adding a replica is instant and can be done during an incident

context