skip to content

When contention is severe enough that even reserved capacity can't keep the high-priority queue's oldest messages within their target latency, what design options exist to protect the high-priority tier further, and what do they cost?

level: seniorimportance: should knowfreq 45%

answer

  1. scheduling changes order; admission control changes total load
  2. producer-side rejection / 429 + Retry-After
  3. load shedding only for value-decays-with-time or droppable work
  4. dynamic capacity shrinkage under load
  5. must be a pre-agreed business decision, not silent

basics

~20 s

If just going first isn't enough anymore, you can also refuse or delay some of the low-priority work entirely — like temporarily rejecting or throttling non-urgent requests — so the important stuff always has enough room. It means some low-priority work gets sacrificed on purpose.

solid answer

~40 s

Beyond ordering, you can apply backpressure/admission control at the producer side: rate-limit or reject low-priority messages before they're even enqueued once system-wide load crosses a threshold, rather than relying purely on consumer-side priority ordering to protect high-priority latency. This can mean returning a 429/'try later' to low-priority producers, dropping the oldest low-priority messages (load shedding) once the queue exceeds a size/age threshold, or dynamically shrinking the resource share available to low-priority consumers even further under sustained overload. The cost is that some low-priority work is deliberately lost or delayed rather than merely deprioritized, so this only makes sense when the business has explicitly decided low-priority work is allowed to fail/degrade under extreme load, and producers/callers need a defined contract (error code, retry-after) for what happens when their message is rejected.

go deeper

for a junior

Should grasp, once explained, that reordering alone can't help if there simply isn't enough total capacity, and that sometimes low-priority work gets rejected on purpose.

for a middle

Should be able to name at least one admission-control technique (e.g., rejecting low-priority requests under load) as distinct from reordering.

for a senior

Should distinguish scheduling from admission control clearly, propose at least two concrete techniques (rejection, load shedding, dynamic capacity shrinkage) with their appropriate use cases, and flag that load shedding is unsafe for correctness-critical work.

for a principal

Should treat this as a business-requirements problem as much as a technical one — driving the decision of which traffic classes are droppable/reject-able through explicit cross-team agreement, and designing the caller-facing contract (error codes, retry semantics, monitoring) that admission control depends on.

## Scheduling versus admission control Priority ordering and reserved capacity, discussed as the standard mitigations for starvation, both assume the system has enough total capacity to eventually get through all traffic — they change the order work is done in, not the total amount of work the system can absorb. Under severe, sustained overload, that assumption breaks: if the volume of incoming high-priority work alone exceeds total processing capacity, no amount of reordering or reserved capacity for low-priority messages can prevent high-priority latency from degrading, because there simply isn't enough throughput to go around even for the high-priority tier by itself. At that point, priority ordering has done everything it can, and further protection for the high-priority tier requires actively reducing the total amount of work entering the system — this is the shift from **scheduling** (deciding order) to **admission control / backpressure** (deciding whether work is accepted at all). ## Producer-side rejection The most direct form is producer-side rejection: before a low-priority message is even enqueued, the producing service checks current system load (queue depth, oldest-message age, a load-shedding flag published by the consumer side) and, if it's past a threshold, declines to enqueue the message at all — returning an explicit error (commonly HTTP `429 Too Many Requests` with a `Retry-After` header, or an equivalent async rejection) to whatever called it. This is strictly cheaper for the system than accepting the message and processing it later, because the message never consumes any queue storage, consumer capacity, or downstream resources at all. The cost is pushed back to the caller, which now needs an explicit contract for what to do when its low-priority work is rejected: - retry later, - drop it, - or surface a user-facing message. And that contract has to be designed deliberately, not left implicit, otherwise callers end up hammering the system with retries that make the overload worse. ## Load shedding inside the queue A second option is load shedding inside the queue itself: once a low-priority queue exceeds a defined size or its oldest message exceeds an age threshold, the system actively drops the oldest (or a sampled subset of) messages rather than continuing to accumulate an ever-larger backlog that will simply take longer and longer to work through even after the surge subsides. This is appropriate only for low-priority work whose value genuinely decays with time or that's acceptable to lose outright — e.g., real-time analytics events where a late data point is worthless, versus something like a billing event that must eventually be processed correctly no matter how late. Applying load shedding to any message with a correctness requirement (must eventually happen, exactly once) is a serious mistake; it's only safe for genuinely best-effort work. ## Dynamic capacity shrinkage A third, softer option is dynamic capacity shrinkage: rather than an all-or-nothing accept/reject decision, the fraction of shared or reserved capacity available to the low-priority tier is reduced further as high-priority load increases — effectively a more aggressive, load-adaptive version of the weighted-polling technique used for ordinary starvation prevention, but with a floor that can shrink toward (though ideally never fully to) zero under the most extreme conditions, then recover automatically once load subsides. ## The cost all three share The unifying cost across all three options is the same: they require the business to have made an explicit decision, ahead of time, that low-priority work is allowed to be: - delayed indefinitely, - rejected outright, - or lost, under defined overload conditions. And that decision needs to be documented and communicated to whatever teams or systems produce that low-priority traffic, since 'sometimes your message just silently never gets processed' is a very different contract than 'your message is guaranteed to eventually process, just possibly slowly.' Skipping this step is a common real failure: a team adds load shedding to solve an incident, and later discovers the 'best-effort' queue they were shedding from was actually being relied on for something the business considered mandatory (e.g., compliance audit logging), producing a second, worse incident. ## Where it shows up A concrete example: a video-streaming platform's transcoding pipeline prioritizes user-uploaded content going live imminently over back-catalog re-encodes for a new codec. During a traffic surge that saturates transcoding capacity even for live-imminent content, the system automatically pauses accepting new back-catalog re-encode jobs at the producer (returning a 429 to the internal scheduler that queues them) once live-imminent queue age crosses two minutes, resuming normal acceptance once the surge passes — a deliberate, pre-agreed trade that the back-catalog work can simply wait rather than compete at all during a true capacity crunch.

  • Why is load shedding dangerous to apply to a queue carrying billing or compliance events?
    Load shedding assumes the dropped work's value decays with time or is acceptable to lose outright, but billing/compliance events typically have a correctness requirement — they must eventually be processed, exactly once, regardless of delay. Dropping them under load doesn't relieve pressure safely; it creates a silent data-loss incident that's often discovered much later, during reconciliation or an audit.
  • What operational contract needs to exist before enabling producer-side 429 rejection for low-priority traffic?
    Callers of the low-priority path need a documented, explicit expectation of what a 429 means for them — retry with backoff, drop the request, or surface a user-facing delay message — decided ahead of time rather than discovered during an incident. Without that, callers commonly retry naively and immediately, which can turn a load-shedding safety mechanism into a retry storm that makes the overload worse.
  • How does dynamic capacity shrinkage differ from the fixed weighted-polling/reserved-capacity technique used for ordinary starvation prevention?
    Weighted polling and reserved capacity use a fixed split (e.g., a permanent 80/20 or a fixed minimum consumer count) regardless of current load. Dynamic capacity shrinkage adjusts that split in response to real-time load, tightening the low-priority tier's share further as high-priority pressure increases and relaxing it again once load subsides, trading some implementation complexity for a system that adapts to the actual severity of contention rather than a static worst case.

It's like a restaurant that, once every table is fully booked and the kitchen is at capacity, stops taking new reservations entirely rather than accepting them and making everyone — including the people with confirmed reservations — wait even longer.

saying these in an interview costs you the question

  • Assumes priority ordering alone can protect high-priority latency under any level of overload
  • Suggests load shedding for correctness-critical (must-eventually-process) work
  • No mention of needing a defined caller contract for rejected/shed low-priority work
  • Treats admission control and scheduling/ordering as the same mechanism
  • Doesn't distinguish 'delayed' from 'lost' when discussing overload mitigations

context