skip to content

An asynchronously invoked AWS Lambda function charges customers, and finance reports occasional double charges even though CloudWatch shows no function errors for those invocations. Explain how a Lambda invocation that succeeded can still run twice, and how you would stop the double charge.

level: seniorimportance: should knowfreq 52%

answer

  1. the guarantee is at-least-once
  2. success and completion are not the same event
  3. the side effect landed, the invocation did not
  4. key off the order, not the attempt
  5. conditional write before the charge

basics

~20 s

Asynchronous Lambda invocation is at-least-once: the internal queue can deliver the same event more than once even without a handler error, and a handler that completes its side effect but then times out is retried too. The fix is idempotency in the handler, not tighter retry settings.

solid answer

~50 s

Asynchronous invocation gives an at-least-once guarantee, not exactly-once, so Lambda may deliver the same event twice on its own — and the second delivery looks like a perfectly clean, error-free invocation in your metrics. Two other paths produce the same symptom without an error appearing: a handler that charges the customer and then times out or dies before returning is counted as a failure and retried, so the charge happens twice while only one invocation is marked failed; and the upstream producer may have published the event twice. Turning retries off is not the answer — it just converts double charges into lost charges. The fix is to make the handler idempotent: derive a key from the business payload, guard the side effect with a conditional write to a durable store, and use the payment provider's own idempotency key so the second call is a no-op rather than a second charge.

code

javascript · 26 lines
javascript
import { DynamoDBClient } from "@aws-sdk/client-dynamodb";
import { DynamoDBDocumentClient, PutCommand } from "@aws-sdk/lib-dynamodb";

const ddb = DynamoDBDocumentClient.from(new DynamoDBClient({}));
const chargeCustomer = async (event) => { /* call the payment API */ };

export const handler = async (event) => {
  try {
    await ddb.send(new PutCommand({
      TableName: "idempotency",
      Item: {
        pk: event.orderId,
        expiresAt: Math.floor(Date.now() / 1000) + 86400,
      },
      ConditionExpression: "attribute_not_exists(pk)",
    }));
  } catch (err) {
    if (err.name === "ConditionalCheckFailedException") {
      return { status: "duplicate-suppressed" };
    }
    throw err;
  }

  await chargeCustomer(event);
  return { status: "charged" };
};

go deeper

for a junior

Know the phrase 'at least once' and what it implies: your handler can see the same event twice, so it must not assume a single execution per event.

for a middle

Explain the partial-success case concretely — side effect completes, invocation then fails, whole event is retried — and describe a conditional write on a business key as the guard.

for a senior

Show production judgment: reject the retry-disabling fix out loud, choose where the idempotency guarantee should live (ideally the downstream system), handle the marker-crash window, and make the suppression observable.

for a principal

Own the guarantee at system level: decide whether duplicates or losses are the cheaper failure for this domain, push stable event identifiers onto producers as a contract, and set the standard for idempotency across every asynchronous handler in the estate.

## The guarantee you actually have Lambda's asynchronous invocation path is documented as **at least once**. That phrase is the whole answer to the first half of the question. Lambda's internal queue prioritises never losing an event over never repeating one, so under some conditions — an internal retry, a redelivery inside the queueing layer — the same event is handed to your function more than once even though your code never threw. Nothing in your CloudWatch `Errors` metric will hint at it, because from Lambda's point of view both invocations succeeded. ## The three ways one event becomes two side effects **Redelivery.** The queue itself may deliver twice. Rare, but not something you may design against by assuming it away. **Partial success then failure.** This is by far the most common cause in practice, and it is entirely your own code. The handler calls the payment API, the charge lands, and then the invocation dies — a timeout a moment later, an out-of-memory kill, an exception in the audit write after the charge. Lambda sees one failed invocation and retries the *entire* event. Attempt two charges the card again. The metric shows a single error, the customer sees two charges, and nothing about that is inconsistent. **Duplicate production.** The upstream service published the same logical event twice, often for exactly the same reason one layer up. A fourth appears on the synchronous path: a caller whose `Invoke` times out at the network level cannot know whether the function ran. If that caller retries, the work happens twice. A timeout is an *unknown*, never a *no*. ## Why tightening the retry policy is the wrong fix The instinct is to set `MaximumRetryAttempts` to 0. That does reduce duplicates, and it buys them at the price of dropped work: every transient dependency blip now silently loses a charge instead of duplicating it. For money, both outcomes are bad, and the lost one is usually worse because it is invisible. Retry settings are a throughput and cost knob; they are not a correctness mechanism. ## The fix: idempotency at the side effect Make the second execution of the same event harmless. Three layers, in order of preference: **1. Use the downstream system's idempotency support.** Payment providers accept an idempotency key on the charge request and return the original result for a repeat. This is the strongest option because the guarantee lives where the money moves. Derive the key from the business payload — an order ID — not from anything Lambda generates per attempt. **2. Guard with a conditional write.** Before the side effect, write a marker keyed on the business identifier with a condition that it must not already exist. In DynamoDB that is a `PutItem` with `ConditionExpression: attribute_not_exists(pk)`; the second attempt fails with `ConditionalCheckFailedException` and you exit early. ```javascript await ddb.send(new PutCommand({ TableName: "idempotency", Item: { pk: event.orderId }, ConditionExpression: "attribute_not_exists(pk)", })); ``` The subtlety worth stating out loud in an interview: a naive marker-then-act sequence trades duplicate execution for *lost* execution, because a crash after the marker but before the charge means every retry now exits early. Production implementations store a state on the marker — in progress, completed, with the result — and let an in-progress record expire so a genuinely crashed attempt can be retried. Powertools for AWS Lambda ships an idempotency utility that implements exactly this over DynamoDB, and reaching for it is a better answer than hand-rolling the state machine. **3. Design the operation to be naturally idempotent.** A write that sets a value is safe to repeat; one that increments is not. Where you can restate the operation as "make it so" rather than "do it again", you no longer need a ledger. ## Choosing the key The key must identify the *business event*, not the attempt. Anything Lambda mints per invocation is useless as a dedupe key, since the whole point is that the two executions are different invocations of the same logical event. If the producer does not supply a stable identifier, that is the actual bug — fix it at the producer rather than hashing the whole payload and hoping. ## Proving it in an interview A strong answer ends operationally: add a metric for suppressed duplicates so you can see the mechanism working, alarm on it if the rate jumps (a spike means something upstream is misbehaving), and keep an on-failure destination so events that genuinely exhausted their retries are captured for deliberate replay — replay being safe precisely because the handler is now idempotent.

  • Why is setting MaximumRetryAttempts to 0 a poor fix for duplicate charges?
    It converts a visible failure mode into an invisible one. Retries exist because most asynchronous failures are transient; removing them means a one-second dependency blip permanently loses the charge, with no error surfacing to anyone. You still have not addressed at-least-once redelivery, which happens without any retry at all.
  • What is the flaw in simply writing an 'already processed' marker before doing the work?
    If the invocation dies between the marker and the side effect, every retry sees the marker and exits early, so the work never happens. Real implementations record a state — in progress versus completed, plus the stored result — and expire in-progress records so a crashed attempt can be retried rather than permanently suppressed.
  • Does the same duplicate-execution risk exist on the synchronous invocation path?
    Yes, but the retry lives in the caller. A caller whose Invoke times out at the network layer cannot tell whether the function ran; if it retries, the work happens twice. A timeout is an unknown outcome, not a failure, so callers that retry timeouts need the same idempotency guarantees from the handler.
  • How would you make the duplicate suppression observable?
    Emit a custom metric each time the idempotency guard short-circuits an execution, and log the business key. A low steady rate confirms the mechanism is working; a sudden spike usually means an upstream producer started double-publishing, which is a defect you want to see immediately rather than absorb silently.

saying these in an interview costs you the question

  • Claims Lambda guarantees exactly-once asynchronous delivery
  • Fixes duplicates by disabling retries entirely
  • Uses a per-invocation identifier as the deduplication key
  • Assumes no Errors metric means no invocation was repeated
  • Believes a client timeout proves the function did not run

context