During a traffic surge your rented broker refuses to create another stream at a ceiling no request lifts — what can you do that hour?
answer
- the hour, not the ticket
- relief differs by ceiling class
- moves on your side of the boundary
- consolidate, re-key, split, retire
- history ceiling has no relief
basics
~20 sOnly moves that need no action from the provider land inside an hour: carry the new channel on an existing stream behind a routing attribute, re-key so fewer objects carry the work, or shed a stream you can retire now. Everything else is a project.
solid answer
~50 sA ceiling nothing lifts cannot be answered with a request, so inside the incident you are limited to moves entirely on your side. The three that fit an hour are: **carry the new channel on an existing stream**, separating it by a routing attribute readers filter on; **re-key or consolidate**, so the same work travels over fewer objects; and **retire something**, if a stream is genuinely dead and its readers are gone. The moves that do not fit are the ones people reach for first — standing up a second cluster and splitting the estate, or migrating to a differently-built tier, both of which are days of work. And one class has no relief at all inside the incident: if the tier never kept history that far back, there is nothing to read, whatever you do now.
go deeper
The point to take away is that some ceilings on a rented broker are not negotiable at all, so asking the provider is not a plan. Learn to recognise a refused create call as a design problem rather than as an outage.
Explain which moves need no provider action and why that is the test for what fits inside an incident. Be able to say what consolidating channels onto one stream actually costs the readers afterwards.
Demonstrate that you triage by ceiling class first, name the one class with no relief before spending the hour on it, and separate the workaround you took from the redesign it defers.
The call you own is which ceiling the estate is allowed to approach. Choosing between governing objects tightly and splitting the estate is a staffing and ownership decision as much as a technical one, and it has to be made before the surge, not during it.
## Why the ceiling surfaced now A hard cap on a hosted broker is a ceiling the provider will not raise however you ask. Normal traffic sits well below one, so the design never exercises it. What reaches it is precisely the unusual event — a surge that spawns per-tenant channels, a replay that opens hundreds of connections, an incident response that creates a parallel stream to divert traffic onto. The cap is therefore met in the worst possible context: under load, with attention elsewhere, by the very action intended to relieve the problem. The first discipline is to stop treating it as a supply problem. Nothing you file changes the number today. The only question worth the next five minutes is which class of ceiling you are against, because the available moves are different for each. ## What the hour actually permits | ceiling class met | move that can land inside the incident | move that cannot | |---|---|---| | object counts (streams, queues, partitions) | carry the new channel on an existing stream behind a routing attribute; consolidate two low-volume channels; retire a genuinely dead stream | creating a second cluster and rebalancing the estate onto it | | connections and sessions | reduce how many connections the fleet opens — fewer instances, shared connections per process, stopping non-essential readers | re-architecting the client fleet's connection model | | single-message size | split one oversized message into several the reader reassembles, or trim fields the consumer does not use | changing how the domain represents the object | | outstanding work per subscription | acknowledge faster by shedding slow work, or add consumers so fewer messages sit outstanding at once | changing the acknowledgement design | | kept history | nothing — see below | nothing | Three of those rows share a property worth naming explicitly: they are moves **entirely on your side of the boundary**. They need no provider action, no ticket and no deployment of anything the provider owns. That is the test to apply to any suggestion made in the incident channel — if the move depends on somebody at the provider doing something, it is not an option this hour, however reasonable it sounds. ## The class with no relief The history ceiling is different in kind. Object counts, connections and outstanding work all describe things that exist right now and can be rearranged right now. The longest history a tier will keep at all describes something that has already happened: records older than that ceiling were never retained, so no action inside the incident brings them back. If the plan was "we will read from the beginning and rebuild", and the tier's ceiling on kept history is shorter than the window you need, the plan was never available. That is worth saying plainly during the incident rather than after an hour of investigation, because it redirects the response to a different source of truth. ## Why it is a redesign, not a ticket Once the incident is over, the distinction that matters is between a **workaround** and the **redesign** it is standing in for. The moves above buy time and each carries a cost: - Consolidating channels onto one stream makes every reader filter, which spends reading capacity on messages the reader discards and couples teams that were separate. - Re-keying changes how work is distributed, which on platforms that split a stream changes ordering relationships between records that used to travel together — a semantic change made under pressure. - Splitting one message into several makes the reader responsible for reassembly and for the case where one part never arrives. - Shedding readers to free connections leaves work unread, which becomes the next incident. So the follow-up work is real: decide which ceiling the estate will be designed away from, and do that deliberately. That is usually either a governance change (fewer, larger, shared streams with an owner per channel, instead of one per tenant) or an estate split (a second cluster with a defined division of which streams live where, and an asynchronous copy where both need the data). Both are weeks, which is exactly why neither is an answer at the moment the ceiling is met. ## Preventing the rediscovery The reason this question is asked at senior level is that the prevention is unglamorous and almost never done: measure your current count in each ceiling class as an ordinary operating signal, next to the ceiling, and treat the ratio as capacity like any other. A ceiling you are at eighty per cent of is not headroom — the same surge that makes you want another stream is the one that consumes the rest.
- What is the cost of consolidating several channels onto one stream under pressure?Every reader now receives messages it must discard, which spends read capacity and can push readers behind. It also couples teams that were independent: a change to one channel's shape now lands in everyone's path, and permissions that were per stream become coarser than they were.
- How do you stop meeting the same ceiling six months later?Track your current count in each ceiling class as an operating signal next to the ceiling itself, and treat a high ratio as consumed capacity rather than as spare room. Then choose deliberately whether the estate grows by governing objects more tightly or by adding a second cluster with a defined division.
- Why is standing up a second cluster the wrong answer inside the incident even when it is the right answer overall?It changes where clients connect, how permissions are granted and which streams live where, and none of that is safe to decide under load. It is the correct destination and a multi-week journey; proposing it during the incident usually just delays the move that would have worked in the hour.
saying these in an interview costs you the question
- Answers 'open a ticket' for a ceiling the provider will not raise
- Thinks a second cluster can be stood up inside an incident
- Believes shortening kept history frees stream-count headroom
- Claims every ceiling class has an hour-scale workaround
- Treats consolidating channels as free rather than as coupling
- Sheds readers to free connections and calls the incident resolved