As a platform lead, how would you set one default partition count for every new stream across a large estate?
answer
- a default is a policy, not a number
- the two errors are asymmetric
- cheap default, published tiers
- make the raise routine
- own the estate-wide unit total
basics
~20 sSet a small default for the long tail of low-rate streams and publish one or two higher tiers a team opts into with a stated reader-parallelism need. A generous single default is multiplied by every stream anyone ever creates.
solid answer
~50 sTreat it as a policy with two opposite failure modes rather than a number. A default that is too low means teams meet the ceiling and need a scheduled raise; a default that is too high is multiplied by copies and by thousands of streams, and pushes up the cluster's total unit count, its metadata set and how long it takes to recover a lost node. Since raising is a supported operation and shrinking is not, the asymmetry argues for a **modest** default plus a few documented tiers a team opts into by stating the reader parallelism it expects. Make the raise a normal, supported request rather than an exception, so nobody over-provisions defensively. Then watch the estate-wide total of parts times copies as a capacity number in its own right, because that is the number a generous default actually moves.
go deeper
Recall that new streams get a default number of parts, and that the default matters because most streams will never have it changed by anyone.
Explain the two failure modes: too low costs a scheduled raise for one team, too high costs standing overhead multiplied across the whole estate.
Argue from the asymmetry - a raise is supported, a shrink is a migration - and show what you would measure on the cluster before setting the level.
Own the policy and what it forecloses: tiers with a stated justification, a routine raise so nobody over-provisions defensively, and a named owner for the estate-wide unit total.
A per-stream number is an engineering decision. A default applied to every stream an organisation will ever create is a policy, it is multiplied by thousands, and it will outlive everyone who argued about it. ## The two failure modes are not symmetric | default too low | default too high | |---|---| | teams meet the ceiling and cannot scale readers | every stream carries unusable parts forever | | the fix is a supported raise, planned around key placement | there is no supported way down; the fix is a migration per stream | | the pain is local, visible and lands on one team | the pain is diffuse, invisible and lands on the cluster | | discovered quickly, when someone tries to scale | discovered slowly, as recovery times and metadata grow | That table is the whole argument. A low default fails loudly, locally and reversibly. A high default fails quietly, globally and expensively. When two errors are this asymmetric, aim at the recoverable one. ## What a defensible policy looks like 1. **A small default for the long tail.** Most streams in any estate carry a trickle and have one or two readers. They should cost the minimum, and they should not require anyone to think. 2. **Two or three published tiers above it.** A team opts into a higher tier by stating the reader parallelism it expects and the interval it expects it in. The point of making them state it is not gatekeeping; it is that a number with a reason attached can be revisited, and a number without one cannot. 3. **A supported raise path.** Publish that increasing the count is a normal request with a normal turnaround. If a raise feels like an exception, every team will over-provision at creation to avoid ever asking, and the policy's cheap default will be quietly defeated. 4. **An estate budget.** Track the total of parts times copies across the cluster, alongside bytes and rates, and set a level at which the tiers get revisited. That total is what a generous default actually moves. 5. **The reasoning stored next to the stream.** Whoever inherits it needs to know what it was sized for, or they will neither trust it nor touch it. ## Why a single universal number is the wrong answer A uniform count is attractive because it needs no conversation, and it is wrong for exactly the reason a uniform machine size is wrong: the estate is not uniform. A handful of streams carry most of the traffic and need real parallelism; the great majority carry very little. One number either starves the first group or taxes the second, and in practice it does both. Tiers are what a uniform number becomes once you admit the distribution. ## What the policy forecloses This is the part a principal is expected to say out loud. A default is inherited silently by every stream created after it, including thousands nobody will ever revisit. If it is generous, the organisation is committed to that overhead for the life of those streams, because there is no supported way to come back down and no team will run a migration for a cost they cannot see on their own dashboard. If it is stingy and the raise path is undocumented, the organisation is committed to a recurring ticket that every team learns to route around by asking for a large tier they do not need. Either way the policy, not any individual choice, is what the estate ends up shaped by. ## Where platforms differ, and why it matters here How much an individual part costs is not the same everywhere: some platforms carry per-part state on every node holding a copy, some separate serving from storage so the cost lands differently, and a rented offering may price or cap the count rather than exposing its cost at all. So the *level* of the default has to be set against your own platform and measured on your own cluster. What travels between platforms is the shape of the policy: a cheap default, published tiers with a stated reason, a raise that is routine, and an estate-wide total that somebody owns. ## The answer that lands Name the asymmetry, refuse the single universal number, describe tiers with a stated reader-parallelism justification, insist that the raise be routine so defensive over-provisioning has no motive, and name the estate-wide unit total as the metric the policy is judged by. Do not quote a specific number as though it were universal - the interviewer is testing whether you know what the number is made of.
- Why does making the raise path routine matter more than the default's exact level?Because the alternative is defensive over-provisioning. If teams believe a raise is a hard-won exception, they will ask for a large count at creation whatever the default says, and the estate inherits the high-default failure mode anyway. A routine raise is what makes a cheap default survivable and therefore real.
- What would make you revisit the tiers?The estate-wide total of parts times copies crossing the level you set, recovery after a lost node taking materially longer than it used to, or a pattern in raise requests showing that most teams are outgrowing a tier within months. All three say the distribution has moved and the tiers were fitted to the old one.
- Should the default differ between streams and shared queues?There is nothing to default for a shared queue - it has no structural count, and its useful reader number is discovered from the work and the downstream resources. The policy therefore applies only to streams that are split into parts, and saying so prevents teams from looking for a setting that does not exist.
saying these in an interview costs you the question
- Answers with one universal number and no reasoning
- Sets a generous default because raising feels risky
- Ignores that the default is multiplied by copies and streams
- Treats a raise as an exception rather than routine
- Nobody owns the estate-wide unit total
- Assumes every stream in the estate has similar traffic