skip to content

Two concurrent invocations of the same stateless Lambda function both read a counter value of 5 from a DynamoDB item, increment it locally to 6, and write 6 back. What went wrong, and what two DynamoDB mechanisms would prevent it?

level: seniorimportance: should knowfreq 55%

answer

  1. lost-update / read-modify-write race
  2. concurrency is normal, not edge case, in serverless
  3. atomic UpdateItem ADD = server-side increment
  4. ConditionExpression = optimistic concurrency check
  5. silent failure, no error thrown

basics

~20 s

Both invocations read the same starting number before either wrote back, so one increment silently overwrites the other and the counter ends up one short. DynamoDB fixes this with atomic updates or conditional writes that check the value hasn't changed.

solid answer

~50 s

This is a classic lost-update race: read(5) by A, read(5) by B, write(6) by A, write(6) by B — the correct result was 7, but B's write blindly overwrote A's, based on a value that was already stale by the time B wrote. It happens because 'read, modify locally, write back' is not atomic across two separate round trips, and serverless amplifies this because concurrent invocations of the same function are the normal, expected case. DynamoDB provides two mechanisms to prevent it: an atomic UpdateItem with an ADD (or SET x = x + :n) expression, which increments the value server-side in one operation so both invocations' increments are applied in sequence rather than racing; and conditional writes (ConditionExpression), which let a write specify 'only succeed if the value is still what I originally read,' failing B's write so the caller can retry with the fresh value instead of silently clobbering A's update.

go deeper

for a junior

Should be able to describe, in plain terms, that both invocations saw the same starting number before either saved their change.

for a middle

Should name 'race condition' or 'lost update' and know that some kind of atomic or conditional operation is needed instead of plain read-then-write.

for a senior

Should explain both DynamoDB mechanisms (atomic update expressions and conditional writes) accurately, including when each is the right tool.

for a principal

Should generalize the pattern to other state stores and to complex business invariants beyond simple counters, and reason about when optimistic concurrency's retry storms become a scalability concern of their own under high contention.

## The timeline of the lost update The scenario is the textbook **lost-update anomaly**, and it's worth being precise about the timeline to see exactly where it goes wrong. 1. Invocation A issues a read and gets back the value 5. 2. Before A writes anything, invocation B — running concurrently, in a completely separate execution environment, because the function is stateless and the platform is free to run as many parallel instances as needed — also issues a read and gets back 5, since A hasn't written yet. 3. A now computes 5+1=6 locally and writes 6 back to the table. 4. B, independently, computes 5+1=6 from its own earlier read and writes 6 back too. 5. The final stored value is 6, even though two increments happened and the correct result should be 7. Crucially, no individual operation was wrong — each read was accurate at the moment it happened, and each write correctly stored what that invocation computed — the bug is that the whole 'read, modify, write' sequence for one invocation isn't atomic with respect to the same sequence in the other invocation; there's a window between A's read and A's write during which B can read the same stale value. ## Why serverless makes this the normal case This matters specifically for stateless serverless architectures because the entire concurrency model is built around exactly this pattern being common and expected, not rare. A traditional single-process application server can sometimes get away with an in-process lock serializing access; a Lambda function has no equivalent, because concurrent invocations of the same function are literally separate, unrelated processes with no shared memory or coordination primitive between them by default — no in-process sharing across instances is precisely why this race is not a corner case here but the normal operating condition under any real load. ## The two mechanisms DynamoDB provides two complementary mechanisms that avoid the race entirely by removing the 'read, modify locally, write back' round trip for the specific case of applying a delta to an existing value. 1. **The first is an atomic update expression.** DynamoDB's `UpdateItem` API supports an `ADD` action or a `SET counter = counter + :incr` expression that is evaluated server-side, atomically, as part of a single request: the server reads the current stored value and applies the increment in one indivisible operation, so two concurrent `UpdateItem` calls with `ADD` are simply serialized by the server, producing the mathematically correct 5→6→7 regardless of the order or timing of the two calls, and critically, the client-side code never needs to read the value first at all. 2. **The second mechanism is optimistic concurrency via conditional writes.** The client reads the value along with a version number or the value itself, then issues a write with a `ConditionExpression` such as 'only apply this write if counter still equals 5.' If another writer's update already changed the stored value before this write arrives, the condition fails, DynamoDB rejects the write with a `ConditionalCheckFailedException`, and the caller is expected to re-read the fresh value and retry the whole operation — this is the right tool when the update logic is more complex than a simple increment, such as 'only reserve this inventory unit if quantity_available > 0,' where an atomic `ADD` alone can't express the business rule. ## The production failure mode The production failure mode when this isn't handled is silent and hard to detect precisely because, as above, every individual read and write looks correct in isolation — there's no error, no exception, no log line, just a counter, an inventory count, or a rate-limit window that's subtly wrong under load, often only noticed much later as a discrepancy during a reconciliation audit or a rate limiter that lets slightly more traffic through than intended under high concurrency. This is exactly why DynamoDB's own well-architected guidance pushes atomic counters and conditional writes as the default pattern for concurrent-safe updates, rather than application-side read-modify-write, and it generalizes beyond DynamoDB: - the same shape of bug and the same fix apply to any externalized state store used from concurrently-invoked stateless functions, - including relational databases via row versioning or `SELECT ... FOR UPDATE` where supported. ## The key takeaway The key takeaway is that concurrency-safety cannot be bolted on after the fact with naive retries around the same racy sequence; it requires either making the operation itself atomic at the store, or making the write conditional on the data being unchanged since it was read, so a stale write fails loudly rather than silently overwriting a concurrent update.

  • Why doesn't simply adding application-level retry logic around the read-modify-write sequence fix this?
    Blindly retrying the same read-modify-write pattern doesn't fix anything by itself, because a retry can hit the exact same race again with a different concurrent invocation; the fix has to change the operation itself to be atomic (server-side ADD) or to detect the staleness explicitly via a conditional write and only then retry with a fresh read.
  • How does a conditional write differ from just re-reading the value right before writing to 'reduce the window'?
    Reducing the window between read and write shrinks the probability of a race but never eliminates it, since there's always some non-zero gap between a read and a subsequent write over the network; a conditional write instead makes the correctness guarantee unconditional by having the database itself atomically check the condition and apply the write in one step, so there is no window at all in which the race can occur.
  • Would using an in-memory global variable in the Lambda function to cache and coordinate the counter instead of DynamoDB fix the race?
    No — as covered by the broader statelessness constraint, concurrent invocations run in separate execution environments with no shared memory, so an in-memory counter wouldn't even be shared between the two racing invocations in the first place; it would make the problem worse, since each instance would maintain its own disconnected count.

Like two people both looking at a whiteboard tally showing '5', each independently adding 1 in their head and writing '6' back — the board ends at 6 instead of 7, and nobody sees an error, because each person's own math was correct; they just didn't know about each other's edit.

saying these in an interview costs you the question

  • Proposes fixing the race by 'just retrying' without making the underlying operation atomic or conditional
  • Suggests an in-memory variable or lock as a coordination mechanism across concurrent Lambda invocations
  • Doesn't recognize read-then-write-back as inherently non-atomic across two network round trips
  • Treats the race as a rare edge case rather than an expected outcome of normal serverless concurrency
  • Confuses DynamoDB's read consistency setting (eventually vs strongly consistent reads) with concurrency-safe writes — they solve different problems

context