A single shared queue has no partition count to pick, so what decides how many competing readers it can usefully be planned for?
answer
- the queue is not the constraint
- arrival rate times work per record
- add a spare share
- waiting moves, it does not vanish
- serialised work never spreads
basics
~20 sThe slowest resource every reader shares while handling a record - a database, a pool, an external service's allowed call rate. The queue imposes no count, so the useful number of competing readers is the one that saturates that resource and no more.
solid answer
~50 sPlan it from the work, not from the queue. Because a queue-shaped broker hands each record to whichever reader asks next, the broker will happily accept as many readers as you start; the ceiling sits in what they all touch. The arithmetic is direct: readers needed is roughly arrival rate in records per second multiplied by the seconds of work each record takes, plus a spare share so a slow interval or a dead reader does not put you behind. Then check that number against the shared resource - connection pool slots, contention on the rows being updated, an external service's allowed call rate. If the arithmetic asks for more readers than that resource can serve, more readers only move the waiting from the queue to the resource, usually adding contention on the way. Work that must be handled one at a time per entity does not spread at all, whatever the queue permits.
go deeper
Recall that a shared queue does not limit how many readers you can attach, so the real limit comes from whatever the readers use while handling each record.
Do the arithmetic: arrival rate multiplied by work per record, plus a spare share, then check it against the pool, database or external service every reader touches.
Show the diagnosis under load - completion rate flat while per-record time climbs means the waiting has moved downstream - and pick the lever that is not 'more readers'.
Own the consequence: a queue whose ceiling nobody measured becomes an organisation-wide assumption that scaling readers is free, and that assumption is discovered during an incident.
On a split stream, the parallelism ceiling is a number someone chose. On a shared queue nobody chose one - which does not mean there is no ceiling, only that you have to go and find it. ## Start with the arithmetic The useful reader count follows from two measurements and nothing else: > readers ~ (records arriving per second) x (seconds of work per record) If 200 records arrive each second and each takes 50 ms of end-to-end work, you need about ten readers to keep level. Twenty gets the same work done with each reader idle half the time; five falls permanently behind. To that base add a **spare share** - room so that a busier interval, a slow dependency or the loss of a reader does not immediately put you behind. Two mistakes hide in that formula, and both are worth naming out loud: - **"Seconds of work per record" is wall-clock, not processor time.** Most record handling is waiting on something else, so the number is usually larger than people guess. - **The arrival rate to size for is the busiest sustained interval**, not the daily average. ## Then find the real ceiling The arithmetic gives a number you *want*. The shared resources behind the readers decide the number you can actually *use*: - **A database.** Connection pool slots are finite, and contention on the same rows rises with concurrency, so past some point each reader gets slower exactly as you add readers. - **An external service.** An allowed call rate is a hard ceiling; readers beyond it collect rejections and retries, which is work that looks like throughput and is not. - **A licence-limited or single-threaded component** anywhere in the path. - **The broker itself, mildly.** Every reader is a connection and a stream of delivery and acknowledgement traffic. This is rarely the binding constraint, but it stops the reader count from being genuinely free. When the arithmetic asks for more readers than the narrowest of these supports, extra readers do not add throughput. They relocate the waiting - out of the queue, where it was visible and measurable, and into the resource, where it shows up as contention, pool timeouts and retries. ## The comparison that makes the point | | split stream | shared queue | |---|---|---| | who set the ceiling | whoever created the stream | nobody - it is discovered | | symptom of exceeding it | extra readers hold nothing and idle | extra readers work, and everything gets slower | | how you find it | read the part count | measure the shared resource under load | | how you change it | a structural change to the stream | make the downstream resource faster, or do less per record | The second row is why the queue shape is arguably the more dangerous of the two. An idle reader is obvious. A reader that is working while contributing nothing net looks exactly like a reader that is helping. ## What to do when the ceiling is too low The lever is never "more readers". It is one of: 1. **Reduce the work per record** - fewer round trips, batch the downstream call, cache what is stable. 2. **Raise what the shared resource can serve** - more pool slots where the database can take them, a higher allowed call rate, an index that removes the contention. 3. **Split the work onto separate queues by cost**, so that slow records do not hold readers that short records could be using. 4. **Accept the rate** and make the arrival side shed or defer, rather than pretending readers will absorb it. ## One constraint the queue cannot help with If records for the same entity must be handled one at a time, no reader count spreads that work - the serialisation is in the requirement, not in the broker. Planning a reader count without knowing whether such a constraint exists produces a number that is right on paper and wrong in production. ## Where designs differ Some queue-shaped brokers let a single queue be served by very many concurrent readers with no practical broker-side cap; others become less efficient well before that, because every reader adds coordination cost to the hand-out. Ask what the platform does before promising a number, and in either case verify by measurement rather than by reading a ceiling off a specification.
- How would you tell that you have already passed the useful reader count?Total completion rate stops rising while per-record handling time rises roughly in step with each reader added, and the waiting shows up downstream: pool checkout time, lock waits, or rejections and retries from an external service. The queue itself looks healthier, because the waiting has simply moved to somewhere less visible.
- Does splitting one shared queue into several change the ceiling?Not by itself - if the readers still contend for the same downstream resource, the total is unchanged. Splitting helps when it separates different work: expensive records stop occupying readers that cheap ones could use, and each queue can be given a reader count sized to its own work per record.
- Why is arrival rate alone not enough to size the reader count?Because the same arrival rate needs one reader or fifty depending on how long each record takes to handle. Work per record is the multiplier, it is wall-clock rather than processor time, and it is the number teams most often underestimate when they size from traffic graphs alone.
saying these in an interview costs you the question
- Assumes a queue's reader count has no practical ceiling
- Sizes readers from arrival rate without work per record
- Uses processor time instead of wall-clock per record
- Adds readers when the database is already contended
- Sizes from a daily average rather than the busiest interval
- Thinks splitting the queue raises a downstream ceiling