skip to content

Partition & Queue Count Planning

The parallelism ceiling picked before launch: the partition count where a stream is split, the consumer count where one queue is shared. Asked because it is cheap to raise and painful to lower.

part ofBroker & streaming operationsoverview, primer and where to startread it →
on this pageshow

questions

5

A stream is split into a fixed number of partitions and a shared queue is not split at all - what caps reader parallelism in each?

level: juniorimportance: must knowfreq 70%

answer

  1. two shapes, one number
  2. parts are handed out whole
  3. readers past the count idle
  4. a queue's bound is downstream

basics

~20 s

On platforms that split a stream into partitions, the count is the parallelism ceiling: each partition is normally served to one reader at a time, so extra readers idle. A shared queue has no such count; competing readers simply take more work.

solid answer

~50 s

These are two different shapes, and only one of them makes you pick a number. Where a stream is split, the partition count is fixed at creation and each partition is usually handed to one of the cooperating readers at a time - so the count is a hard ceiling on how many readers can make progress at once, and reader number `n+1` on an `n`-partition stream gets nothing but standby duty. A shared queue has no structural count: every record goes to whichever connected reader asks next, so five readers or fifty all get work. That does not make a queue's parallelism unbounded - it means the bound is outside the broker, in whatever each reader touches while handling a record. Knowing which shape you are on tells you whether the ceiling was a design-time decision or is an operational dial.

go deeper

for a junior

Recall that a stream is split into parts to get parallel reading, and that the number of parts caps how many readers can work at once. A shared queue has no such number.

for a middle

Explain the assignment rule that creates the ceiling - a part is handed to one reader at a time - and say where a shared queue's ceiling really lives instead, which is outside the broker.

for a senior

Show the diagnosis: throughput flat after scaling readers means check the part count before profiling code, and on a queue means find the shared resource behind the readers.

for a principal

Frame it as which shape the platform commits the organisation to: a number every team must choose and defend, against a number nobody picks and everybody discovers downstream.

Parallel reading is not one mechanism. Brokers in this class come in two shapes, and the shape decides whether there is a number to pick at all - which is exactly why this gets asked before a stream exists rather than after. ## Shape one: a stream split into parts Platforms that split a stream divide it into a fixed number of **partitions** when the stream is created. Each partition is an independent part of the stream with its own stored records, and the common cooperating-reader arrangement hands **each partition to one reader at a time**. Two consequences follow at once: - Readers up to the partition count each receive a share of the partitions and do work. - Readers beyond the partition count receive no partitions. They are warm standbys - genuinely useful when one of the others dies, worth nothing for throughput. So the count chosen up front is a **parallelism ceiling**: a cap on how many reader processes can be making progress at once on that stream, fixed before a single record was written. Adding machines to the reading side stops helping precisely there, and no amount of reader-side tuning moves it. One honest qualification: some split-stream platforms also offer a mode in which several readers share the records of a single part. That mode buys parallelism back and gives up the one-reader-per-part assignment, so it is a different trade rather than an exception to the arithmetic - and where it exists you should say which mode you mean. ## Shape two: one shared queue A queue-shaped broker has no structural count to choose. There is one work list, and every record goes to whichever connected reader asks for work next. Start five readers and five get work; start fifty and fifty do. Nothing inside the queue caps that. That does not make the parallelism unbounded. It means the bound is **not in the broker**: it sits in whatever each reader touches while handling a record - a database and its row contention, a connection pool, an external service's allowed call rate, a file system. Past that point, extra readers convert waiting in the queue into waiting at the shared resource, usually with contention costs on top. ## The two shapes side by side | | split stream | shared queue | |---|---|---| | unit handed out | a partition, held by one reader at a time | a single record, to whichever reader asks next | | number chosen up front | yes - the partition count | none | | readers beyond the ceiling | receive nothing; standbys only | still receive work until something downstream saturates | | where the ceiling lives | inside the stream's own shape | outside the broker, in the shared resources readers touch | | moving the ceiling | a structural change to the stream | start or stop reader processes | ## Why this is a first-screen question It separates candidates who have only added readers from candidates who know why adding readers sometimes does nothing. The failure it predicts is concrete and common: a team scales the reading side of a split stream from four processes to twelve, watches throughput stay flat, and starts profiling application code - when the stream was created with four parts and eight processes are holding nothing. On a queue the same team scales to twelve, throughput does rise, and then flattens for an entirely different reason that lives in the database behind the readers. The practical habits that follow: 1. Before scaling readers on a split stream, check the partition count first. It is the cheapest possible diagnosis. 2. Before scaling readers on a shared queue, find the shared resource each record touches and measure how much of it is left. 3. When writing down capacity for a new stream, record which shape you are on - the number means something different in each. ## What this question is not about It is not about ordering, keys or how a record is routed to a particular part; that is its own subject and it is not needed to answer this. It is also not about how you would execute a change to the count on a cluster that is already running. Here the question is only: for each shape, what sets the number of readers that can usefully work at once, and is that number something you chose?

  • If eight reader processes are running against a stream split into four parts, what are the four idle processes worth?
    Availability, not throughput. They are warm standbys: when one of the four working readers dies or is restarted, the work it held is handed to a process that is already connected and ready, which shortens the gap. They add no capacity while everything is healthy, and they still cost connections and machines.
  • Does the writing side face the same ceiling?
    No. Writers are not assigned parts exclusively - many writers can write to the same stream and to the same part concurrently. The partition count caps how many readers can work in parallel, and it also spreads the write load across record-serving nodes, but it does not cap how many writer processes you may run.

A shop with six checkout lanes can keep six cashiers busy; a seventh cashier has no lane and stands about. A shop with one snaking queue feeding whatever counter is free can staff as many counters as the floor space and card terminals support - the queue never sets the number.

saying these in an interview costs you the question

  • Thinks more reader processes always means more throughput
  • Believes a shared queue also has a fixed count to pick
  • Says idle readers are the broker throttling them
  • Assumes a queue's parallelism is genuinely unbounded
  • Confuses number of connections with number of working readers
open as a page

Why is a stream's partition count treated as a one-way door, when raising it later is routine and lowering it is not?

level: middleimportance: must knowfreq 62%

basics

~20 s

The change is asymmetric. Raising the partition count on a stream is a supported, routine operation; lowering it generally is not offered at all, so coming down means creating a second stream at the smaller count and moving writers and readers across.

open as a page

If raising a stream's partition count is routine, why is picking a very large number up front still a mistake?

level: middleimportance: should knowfreq 58%

basics

~20 s

Each partition carries fixed overhead on every node holding a copy of it, multiplied by the copy count and by every stream in the estate. A very large total also lengthens recovery and enlarges the metadata the coordination membership must agree on.

open as a page

A single shared queue has no partition count to pick, so what decides how many competing readers it can usefully be planned for?

level: seniorimportance: should knowfreq 48%

basics

~20 s

The slowest resource every reader shares while handling a record - a database, a pool, an external service's allowed call rate. The queue imposes no count, so the useful number of competing readers is the one that saturates that resource and no more.

open as a page

As a platform lead, how would you set one default partition count for every new stream across a large estate?

level: principalimportance: should knowfreq 42%

basics

~20 s

Set a small default for the long tail of low-rate streams and publish one or two higher tiers a team opts into with a stated reader-parallelism need. A generous single default is multiplied by every stream anyone ever creates.

open as a page