skip to content

A Lambda function is triggered by s3:ObjectCreated:* on a bucket and writes its processed output back into that same bucket. What goes wrong, and how do you design around it?

level: seniorimportance: should knowfreq 40%

answer

  1. the output triggers the trigger
  2. growth is bounded by nobody
  3. the blast radius is not just cost
  4. fence input from output structurally
  5. reserved concurrency as the kill switch

basics

~20 s

The output write fires the same notification, so the function invokes itself in an unbounded loop that burns concurrency, S3 request charges and Lambda duration until someone stops it. Fix it by writing output to a different bucket, or by fencing input and output with disjoint prefix or suffix filters.

solid answer

~50 s

Every object the function writes matches its own trigger, so each invocation creates the next one and the pipeline recurses until you intervene. The damage is not just the bill: the runaway consumes the account's shared Lambda concurrency, so unrelated functions start getting throttled, and S3 request charges accumulate alongside the duration charges. The clean fix is a separate output bucket — a structural fence nobody can accidentally undo. If output must live in the same bucket, make the trigger's filter exclude it: trigger on `prefix: incoming/` and write to `processed/`, or trigger on a suffix the output never carries. As a blast-radius cap, set reserved concurrency on the function so a loop is bounded rather than unbounded, and alarm on invocation count. To stop one that is already running, set reserved concurrency to zero or delete the notification configuration.

code

bash · 9 lines
bash
# Emergency stop: throttle every invocation of the runaway function.
aws lambda put-function-concurrency \
  --function-name Processor \
  --reserved-concurrent-executions 0

# Then cut the source of new events.
aws s3api put-bucket-notification-configuration \
  --bucket processing-bucket \
  --notification-configuration '{}'

go deeper

for a junior

Understand that a write into a bucket is itself an ObjectCreated event, so a function that writes into the bucket it listens to will trigger itself again.

for a middle

Explain the fence: how a prefix or suffix filter on the notification makes the output invisible to the trigger, and why a separate output bucket is sturdier than either.

for a senior

Bring in the blast radius — shared regional Lambda concurrency starving unrelated functions — and give the containment sequence plus the invocation-count alarm that detects a loop in which every invocation succeeds.

for a principal

Own the guardrail rather than the incident: reserved concurrency as a default on event-driven functions, budget and anomaly alerting, and a review rule that every S3-triggered function states why its writes cannot match its own trigger.

## The mechanism An S3 notification does not care who wrote the object. If the function's trigger is `s3:ObjectCreated:*` on the whole bucket and the function does a `PutObject` back into that bucket, the put produces exactly the event the trigger is watching for. Invocation *n* creates the object that causes invocation *n+1*. If each invocation writes one object the growth is linear and merely expensive; if it writes several — one thumbnail per size, say — the growth is exponential and the meter moves very fast. ``` upload -> event -> lambda -> writes output/ -> event -> lambda -> writes output/ -> ... ``` ## What it actually costs Four separate meters run at once, and the third is the one that turns an expensive mistake into an incident: - **Lambda duration and requests** for every invocation in the chain. - **S3 request charges** for every PUT and every GET the function makes, which on a high-rate loop can exceed the compute cost. - **Account-wide concurrency.** Lambda concurrency is a shared, regional pool. A runaway that climbs to the account limit throttles *other* functions — the API-backing function, the queue consumer — and a storage mistake becomes a site outage. - **Downstream amplification.** Whatever the function calls (a database, a third-party API) receives the same runaway rate. ## Fixing it structurally Ranked from most to least robust: 1. **A separate output bucket.** The input bucket has the notification; the output bucket has none. There is no filter to get wrong and no way for a later change to a key convention to reintroduce the loop. This is the default worth arguing for. 2. **Disjoint prefixes with a filter on the trigger.** Trigger on `prefix: incoming/`, write to `processed/`. Correct, cheap, and the standard single-bucket pattern — but it depends on the filter staying in place. If someone later broadens the notification to the whole bucket, or a new code path writes back into `incoming/`, the loop returns. 3. **Disjoint suffixes.** Trigger on `suffix: .raw.json` and write `.out.json`. Same tradeoff as prefixes, and more fragile because extensions are easy to change. 4. **A guard inside the handler.** Check the key against the output convention and return early. This still costs one invocation per written object — it bounds the blast radius but does not remove the recursion, and it is the weakest of the four because the fence lives in code rather than in configuration. A useful discipline for reviewing any S3-triggered function: state, in one sentence, why the objects this function writes cannot match its own trigger. If the sentence is hard to write, the fence is not real. ## Bounding the damage before it happens Even with a correct fence, treat unbounded invocation as a failure mode worth capping: - **Reserved concurrency** on the function sets a hard ceiling on how many copies can run at once. It caps the loop's rate and, just as importantly, protects the rest of the account's functions from being starved by this one. - **A CloudWatch alarm on the function's invocation count** over a short period catches the runaway in minutes rather than at the end of the billing cycle. Duration and error-rate alarms will not: in a loop every invocation succeeds. ## Stopping one that is already running Speed matters more than elegance. Two effective moves: - **Set the function's reserved concurrency to zero.** This throttles every invocation immediately and is fully reversible. It is the fastest kill switch, at the cost of stopping legitimate work too. - **Remove the notification configuration from the bucket.** This stops new events at the source. Anything already in flight still runs. Deleting the objects the loop created is cleanup, not containment — do it after the trigger is severed, or you are deleting into a running loop. And when you re-enable, remember that S3 does not backfill: the legitimate uploads that arrived while the trigger was off will not be replayed, so plan a batch pass over them. ## The interview version Name the loop, name the account-wide concurrency blast radius (this is the detail that separates a senior answer from a textbook one), give the separate-bucket fix as the structural default, and mention reserved concurrency plus an invocation-count alarm as the cap. Then say how you would stop one at three in the morning.

  • The function must write into the same bucket. What exactly do you configure?
    Give the trigger a prefix filter matching only the input area — prefix incoming/ — and write every output under a prefix the filter cannot match, such as processed/. Add a key check in the handler as a second layer, returning early for anything outside incoming/. Document the invariant next to the notification configuration, because the whole design rests on the filter staying narrow.
  • Why does an account-wide problem come out of one misconfigured function?
    Lambda concurrency is a shared regional pool. A function with no reserved concurrency can climb until it consumes what is available, and every other function in that account and Region then gets throttled. Setting reserved concurrency on the risky function caps it, and setting it on the critical ones guarantees them a floor. This is the argument for treating concurrency limits as a standard part of a function's definition.
  • Would a CloudWatch alarm on the function's error rate have caught this?
    No. In a recursion every invocation succeeds, so error rate and duration both look healthy while the bill climbs. The signal is volume: an alarm on the invocation count, or on concurrent executions approaching the account limit, catches it in minutes. Cost anomaly detection catches it eventually, but hours later.

It is a microphone pointed at its own speaker: the output is fed straight back into the input, and the only fixes are moving the microphone or cutting the power.

saying these in an interview costs you the question

  • Blames Lambda for retrying instead of the trigger
  • Thinks a try/catch in the handler prevents recursion
  • Assumes the account concurrency limit safely caps it
  • Deletes the created objects before cutting the trigger
  • Relies on an error-rate alarm to detect the loop

context