Should a shared ticket-intake queue be bounded, and what should enqueue report when it is full?
answer
- every in-memory queue has a limit
- who chooses it, you or the machine
- what does the caller learn when it is full
- the worst outcome is a successful-looking drop
- twenty callers now need a rejection path
basics
~20 sBound it whenever arrivals are not limited by something else, because an unbounded in-memory queue converts a slow consumer into a dead process. A bounded queue must then report refusal explicitly on enqueue — never accept-and-discard — so every caller owns what happens to the rejected ticket.
solid answer
~50 sCapacity is an interface decision, not an implementation detail. Unbounded is defensible only when something upstream already limits how many tickets can be in flight; otherwise the queue silently converts "the consumer is slower than the producer" into "the node ran out of memory", a failure that arrives late, far from its cause, and takes unrelated work down with it. Bounding converts that into a small, local, testable event — but only if the refusal is *visible*. The contract I would ship is an enqueue that returns an explicit accepted-or-refused outcome the caller cannot plausibly ignore, with the refusal counted and observable. Accept-and-discard is the one option to rule out: a ticket disappears and nothing in the system records it. The real cost is organisational — every calling team now writes a rejection path — so the capacity must be configurable and the contract documented once, not reinvented per caller.
go deeper
Know that an in-memory queue holding items indefinitely consumes memory that is not free, and that a queue with a maximum size has to do something specific when a new item arrives and there is no room.
Explain the two mechanics behind the choice: a fixed-capacity backing store that must refuse, versus one that reallocates to grow with amortized constant-time appends. Say what each does to memory and to the caller's code.
Argue from the failure mode. Show that unbounded growth surfaces late and far from its cause while a cap produces a local, timestamped, countable event, and insist that refusal is reported rather than hidden behind a successful call.
Own the contract and its blast radius across every calling team: which outcome shape callers cannot ignore, capacity as configuration budgeted in bytes, the observability that proves the cap is doing work, and the cost of twenty teams each writing a rejection path.
## The decision, stated honestly "Unbounded" is not the absence of a limit. It is the choice to let the limit be the machine's memory, discovered at runtime, at the worst possible moment, with no useful diagnostic. Every in-memory queue is bounded; the only question is whether the bound is one you chose and can talk about, or one the operating system enforces by killing the process. So the first question is not "how big?" but **"what limits arrivals?"** If tickets can only be created by a fixed number of upstream agents each holding one in flight, the depth is already bounded by something real, and adding a second, artificial ceiling buys complexity for nothing. If arrivals are driven by anything you do not control — an incident that generates ten thousand tickets in a minute, a retry loop upstream, a fan-out — then the queue has no ceiling and the process will find one for you. ## Why unbounded fails badly rather than gracefully The failure mode is what matters, not its probability. An unbounded queue whose consumer falls behind grows steadily, and nothing about the growth is visible in the queue's own behaviour: enqueue keeps succeeding, dequeue keeps returning tickets, every operation stays O(1). Memory climbs until the process dies — and it takes down every unrelated thing sharing that process, long after the actual cause (a slow downstream dependency, say) started. The signal is maximally delayed and maximally displaced from the cause. A bounded queue turns the same underlying problem into an event with a timestamp, a location and a count: "intake refused 412 tickets between 14:02 and 14:05". One of those is debuggable at three in the morning; the other is a memory graph and a guess. ## The contract on a full enqueue Once bounded, `enqueue` has an outcome it did not have before, and the shape of that outcome is the real design decision: | Contract | What the caller must do | Failure style | |---|---|---| | Return an explicit accepted/refused outcome | Inspect the result and decide | Local, visible, easy to test | | Signal a distinct queue-full condition on the error path | Handle that specific condition | Visible, but easy to swallow generically | | Accept and silently discard | Nothing — and that is the problem | Invisible data loss | The third is the one to rule out. A discarded ticket that reported success is indistinguishable from a served one at the call site, so the loss surfaces days later as a customer asking why nobody replied. If a queue may drop, dropping must be a *stated* outcome that the caller sees and that a counter records — never a side effect of a successful-looking call. Between the first two, prefer whichever the surrounding code cannot ignore by accident. A returned outcome that callers routinely discard is no better than a silent drop; a distinct error condition that everyone catches generically alongside unrelated failures is no better either. This is a judgment about the calling code you actually have, not a universal ranking. ## Picking the number Capacity is policy, so it belongs in configuration, not in a constant. Derive the starting value from the observed high-water mark under the worst burst you have measured, plus headroom for a burst you have not. Budget it in **bytes, not items** — what the ceiling protects is memory, so capacity times record size is the number that matters, and a queue of large records needs a smaller count than the same ceiling would allow for small ones. Then instrument it: current depth, high-water mark and refusal rate. Depth that never approaches the cap means the cap is doing nothing and the real risk lives elsewhere; a refusal rate that is chronically non-zero means the cap is being used as a substitute for capacity planning. ## The organisational cost, which is the actual trade This queue is a shared component. Bounding it does not merely change one implementation; it hands every calling team a new obligation — a rejection path to design, test, monitor and get right. Multiply that by twenty teams and it is a genuine cost, and the honest framing of the decision is: *unbounded concentrates the failure into one catastrophic, unowned place; bounded distributes many small, owned failures across every caller.* The second is almost always the better system and unambiguously the more expensive change. What makes it affordable is refusing to let each team invent its own answer. Ship one refusal contract, document it once, provide the counters, and — where the platform allows — a shared helper for the common handling so that twenty rejection paths are one reviewed path used twenty times. The engineering judgment here is not "bounded is correct"; it is knowing that the cap is the easy half and the rejection contract is the half that decides whether the change actually improves anything.
- How would you pick the actual capacity number?From the measured high-water mark under the worst burst you have observed, plus headroom, expressed as configuration rather than a constant — capacity is policy. Budget it in bytes rather than items, since the ceiling exists to protect memory and record size turns a count into a footprint. Then watch depth, high-water mark and refusal rate to see whether the number is doing any work.
- It is a shared component. What does bounding it cost the teams that call it?Every call site gains a rejection path it did not have, and a rejection path is code someone must design, test and monitor. That is the real trade: unbounded concentrates the failure into one catastrophic place nobody owns, bounded distributes small handled failures across every caller. Make it affordable by shipping one documented refusal contract and shared handling instead of twenty invented ones.
- What would you monitor once the cap is in place?Current depth, high-water mark and refusal rate over time. Depth that never nears the cap means the cap is inert and the risk is somewhere else. A chronically non-zero refusal rate means the ceiling is standing in for capacity planning — the queue is telling you the consumer is under-provisioned, and that is a different fix from raising the number.
- Is a queue that grows on demand a way to avoid the decision?No — it relocates it. Growth doubling keeps enqueue amortized O(1) and postpones the ceiling, but the ceiling is still there and is now the process's memory, discovered as a crash. Growing is a fine implementation choice underneath either contract; it is not an answer to the question of what the caller is told when there is no more room.
saying these in an interview costs you the question
- Unbounded means the queue can never fail
- Just drop the item quietly and return success
- Pick a round capacity number and hard-code it
- Bounding is purely an implementation detail
- A growing queue removes the need for a ceiling
- Callers will notice a dropped item eventually