skip to content

An SQS-triggered Lambda scales out under load and exhausts the connection limit of the RDS database behind it. How do you cap how much of that queue is processed at once, and why is putting reserved concurrency on the function the wrong lever?

level: principalimportance: should knowfreq 44%

answer

  1. cap the reader, not the executor
  2. throttles still cost a receive
  3. the backlog belongs in the queue
  4. ScalingConfig on the mapping
  5. count connections, then divide

basics

~20 s

Set ScalingConfig MaximumConcurrency on the event source mapping, which tells the poller itself not to exceed that many concurrent invocations. Reserved concurrency instead lets the poller keep receiving messages and be throttled, inflating receive counts and pushing healthy messages toward the dead-letter queue.

solid answer

~60 s

Cap it at the source: the event source mapping takes a `ScalingConfig` with `MaximumConcurrency`, which bounds how many concurrent invocations that mapping will drive. The poller reads only as fast as that allows, so the backlog stays in the queue — where a backlog is supposed to sit — and the database sees a ceiling you chose. Reserved concurrency looks equivalent but works one layer too late: it caps the *function*, not the *poller*. The poller still receives messages, the invocations are rejected as throttles, and those messages return to the queue with their receive count already incremented. Lambda does back off when it sees sustained throttling, but the damage is that healthy messages accumulate receives and can cross the redrive threshold into the dead-letter queue having never actually failed. Reserved concurrency remains the right tool for protecting *other* functions from this one; `MaximumConcurrency` is the right tool for pacing *this* consumer. The deeper answer usually pairs it with RDS Proxy so a pool, not the connection count, absorbs the remaining burst.

go deeper

for a junior

Know that a queue-triggered function scales out on its own as the backlog grows, and that this can overwhelm whatever the function talks to downstream.

for a middle

Be able to name MaximumConcurrency on the event source mapping as the setting that bounds a queue consumer, and describe how it differs from a limit applied to the function itself.

for a senior

Explain the throttle path in detail — receive, refuse, return, receive count incremented — and why that dead-letters healthy messages, then size the cap from the downstream resource.

for a principal

Own the whole backpressure story: where the constraint belongs architecturally, the latency and retention contract a cap implies, pooling versus capping, and the blast-radius guard you keep in place regardless.

## The shape of the failure A queue-triggered Lambda scales on backlog: as the queue grows, Lambda adds pollers, and concurrency rises until the backlog clears or the account concurrency limit is reached. That is exactly the behaviour you want in front of a stateless HTTP call and exactly the behaviour that destroys a relational database, because every concurrent execution holds at least one connection and connection slots are a small, fixed resource. A traffic spike that the queue absorbed beautifully is then converted, by your own consumer, into a connection storm. The queue was supposed to be the shock absorber. Uncapped consumer scaling removes the absorption. ## The right lever: cap the poller The event source mapping carries a `ScalingConfig` object with a single field, `MaximumConcurrency`. It bounds the number of concurrent invocations that *this mapping* will drive, and it is enforced by the poller — accepted values run from 2 to 1000. ```bash aws lambda update-event-source-mapping \ --uuid 1a2b3c4d-5e6f-7a8b-9c0d-1e2f3a4b5c6d \ --scaling-config '{"MaximumConcurrency": 20}' ``` Because the constraint lives in the poller, the poller simply reads less. Messages are never received, so their receive counts never move, nothing is redelivered, and the backlog waits in the queue as designed. `ApproximateAgeOfOldestMessage` grows, which is precisely the signal you want to alarm on: it says "we are behind", not "we are failing". ## Why reserved concurrency is the wrong lever here Reserved concurrency sets an upper bound on the function's concurrent executions (and simultaneously reserves that capacity from the shared account pool). It is enforced at **invoke** time, not at read time. The sequence under load is: 1. The poller receives a batch — the message's receive count increments and it is now invisible for the redelivery window. 2. The poller invokes the function. 3. Lambda rejects the invoke because the reserved limit is reached. 4. The batch is not deleted; the messages come back when the redelivery window expires. 5. Repeat — with the receive count one higher every time. Nothing failed, yet every lap moves those messages closer to the queue's `maxReceiveCount` and therefore to the dead-letter queue. You end up dead-lettering healthy messages because your own throttle kept refusing them. Lambda does apply back-off when it detects sustained throttling of a queue mapping, which softens the effect, but the mechanism is still fundamentally "receive, refuse, return" rather than "do not receive". There is a second reason to prefer the mapping: scope. Reserved concurrency is a property of the function. If the same function is also invoked by API Gateway, capping it for the sake of the database throttles your API too. `MaximumConcurrency` constrains only the queue path. ## What reserved concurrency is genuinely for Do not conclude it is useless. It has two jobs this scenario does not need. As a **ceiling**, it stops one runaway function eating the whole account's concurrency and starving everything else — worth keeping as a blast-radius guard even when the mapping is capped. As a **floor**, the same reservation guarantees this function capacity that no other function can take. Set both if you want: reserved concurrency as the safety net for the account, `MaximumConcurrency` as the pacing control for the queue. ## The rest of the design conversation At this level the interviewer wants the reasoning around the number, not just the API call. **Where does the ceiling come from?** Work backwards from the database: usable connections, minus the headroom migrations and human sessions need, divided by connections held per execution. That is your concurrency, and the mapping is where you enforce it. **Should you be capping at all, or pooling?** RDS Proxy multiplexes many client connections onto a small server-side pool, which changes the arithmetic in your favour and survives failover more gracefully. Capping and pooling solve overlapping problems; the honest answer usually includes both, with the cap sized to the proxy's limits rather than the database's. **Batch size is the other throughput knob.** Throughput is roughly concurrency times batch size divided by duration. If a cap of 20 leaves you too slow, a larger batch does more work per connection rather than demanding more connections — provided the function timeout covers it and partial batch failures are reported. **What is the latency contract?** A cap converts a scaling problem into a queueing delay. Somebody has to own the statement "during a spike, messages may wait N minutes", and that has to be checked against message retention, not just against user patience: a backlog that outlives retention is data loss. **How do you know it is working?** Alarm on the age of the oldest message and on queue depth, not on invocation count. If throttles appear on the function, that is the signal you left the cap in the wrong layer.

  • Concretely, how does throttling from reserved concurrency push a healthy message toward the dead-letter queue?
    Each time the poller receives that message, its receive count increments — before the function is even invoked. If the invoke is throttled, the message is not deleted, it returns to the queue, and the cycle repeats with a higher count. Once the count exceeds the queue's redrive threshold it is dead-lettered, despite never having been processed or failed.
  • How would you actually choose the MaximumConcurrency number?
    Derive it from the constrained resource, not from Lambda. Take the database's usable connections, subtract headroom for migrations, replicas and human sessions, and divide by the connections one execution holds. Then sanity-check the resulting throughput against arrival rate and message retention, so the backlog you are choosing to accept still drains before messages expire.
  • Is there still a reason to set reserved concurrency on this function?
    Yes, but for a different purpose: as a blast-radius guard. It stops this function from consuming the account's shared concurrency pool and starving unrelated functions, and its floor semantics reserve capacity nothing else can take. Use it as the account-level safety net while the mapping's MaximumConcurrency does the pacing for the queue path.
  • What would you monitor to confirm the cap is behaving as intended?
    Queue-side signals: the approximate age of the oldest message and the visible message count, which show the backlog draining at the rate you chose. On the function side, throttle count should be at or near zero — throttles reappearing means the constraint is being enforced at invoke time again, which is the failure mode you were trying to leave behind.

saying these in an interview costs you the question

  • Uses reserved concurrency to pace an SQS consumer
  • Claims throttled invocations do not affect the message's receive count
  • Caps a function that other triggers also invoke, throttling them too
  • Adds shards or scales the database instead of bounding the consumer
  • Ignores that a capped consumer can back up past message retention

context