Your managed engine will not scale further and the provider says that ceiling is the product's design - what does that change about your options?
answer
- not every ceiling is the same kind
- who owns the number you hit
- policy ceiling against a product ceiling
- one moves on request, one never does
- product ceiling means reduce, split or leave
basics
~20 sA ceiling that is the product's design cannot be raised by asking or by paying more, so it turns a capacity problem into a migration decision: reduce the workload, split it across several instances, or run the same engine yourself on plain machines.
solid answer
~50 sFirst establish which kind of ceiling you hit, because the two lead to completely different plans. Some ceilings are policy: a number the provider sets per account and will move on request, where the cost of hitting it is the time the request takes. Others are architectural: the tier's largest size class, a cap on stored data per instance, a capability the tier simply does not expose. No request raises those, because the ceiling is how the product is built. Ask the provider in exactly those words and get the answer in writing. If it is architectural you have three real moves: reduce demand so the workload fits again, split it across several instances of the same tier, or take the engine back onto plain compute. Only the third removes the ceiling, and it is priced by how much data has to move.
go deeper
Know that a managed tier carries limits you did not choose, and that some of them can be raised on request while others are fixed by how the product is built.
Explain how you tell the two apart: who owns the number, whether it differs between accounts, whether a documented request path exists, and what the provider will put in writing.
Show that you confirm the ceiling's kind before planning anything, and that a product ceiling leaves exactly three moves - reduce demand, split across instances, or self-run the engine.
Weigh which of the three the organisation can actually sustain, since splitting multiplies operational surface and self-running hands back duties, and set the point at which the choice must be made.
## Two ceilings that feel identical from inside the application When a managed service refuses to go further, the application sees the same thing either way: a provisioning call that fails, a write that will not land, a connection that is not accepted, a size class that cannot be selected. The distinction that decides what you do next is **not visible in that error**. It is a question about who owns the number. A **policy ceiling** is a value the provider sets per account or per tier for its own protection - shielding shared capacity, limiting the blast radius of a mistake, keeping a brand-new account from running up an enormous bill. It is not a property of the engine. Somebody at the provider can change it, and the cost of hitting it is measured in **time**: time to notice, to raise the request, to have it approved, plus whatever the workload does meanwhile. An **architectural ceiling** is a property of the product. The tier publishes a largest size class and there is nothing above it; stored data per instance is capped by how the service is built; a capability the engine has upstream is deliberately not exposed. No request raises it, because there is nothing to raise. The cost of hitting it is an **architecture change**. ## How to tell them apart before you plan anything - Ask the provider in those exact words: *is this a limit you can raise for my account, or is it the product's design?* Vague answers are answers - press until the sentence is unambiguous. - Check whether the number varies between accounts, tiers or regions. A number that differs is administered by somebody; a number that is identical everywhere is usually built in. - Look for a documented request path. Where one exists, the ceiling is administered. - Check what sits above your current size class. If you are already on the largest, there is no headroom left to buy. - Get the answer dated and in writing on the ticket. This sentence will be quoted in a planning argument months later. | | Policy ceiling | Product ceiling | |---|---|---| | Who can change it | provider staff, per account | nobody | | What hitting it costs | waiting time | an architecture change | | Typical evidence | a documented request path; the number differs across accounts | the largest size class is already in use; the capability is absent from the tier | | What it forces | a request, plus a bigger margin next time | reduce, split, or leave the tier | ## What a product ceiling actually forces There are three honest moves, and a candidate who names all three is answering the real question: 1. **Reduce demand so the workload fits again.** Trim retention, move cold data to cheaper storage, take the least valuable load off the hot path, push read traffic somewhere else. This buys months, not a solution, and each mitigation can only be spent once - so record which ones you have already used, or the next capacity conversation starts from a false picture of your headroom. 2. **Split the workload across several instances of the same tier.** This **multiplies** the ceiling rather than removing it. You now pay for routing logic, for work that used to be one query and is now several, for a consistent way to add and rebalance instances, and for several times the operational surface - but you stay on the tier and keep everything it was doing for you. 3. **Run the same engine yourself on plain compute.** This is the only move that removes the ceiling outright, because the engine's real limits are then the machine's, not the product's. It hands back every duty the tier was performing, and its schedule is set by how much data must move. ## The mistakes that make this expensive - Assuming every refusal is negotiable, so nobody plans while the runway burns. - Assuming no refusal is negotiable, so a team migrates away from a number that support would have raised in a day. - Jumping straight to the largest size class. That buys **headroom**, not a different ceiling, and it means the next ceiling arrives with nothing left to manoeuvre with. The interval between today and the largest class is your planning horizon, and it should be measured, not felt. - Starting the migration plan before the provider's answer is in writing. - Reading a temporary refusal - throttling under load, a transient capacity shortage - as a permanent design limit. ## How to answer this in an interview Lead with the distinction, then say how you would confirm it, then say what each answer forces. The interviewer is checking that you do not treat every limit as a wall, do not treat every wall as negotiable, and understand that a product ceiling converts an operational problem into a design decision with a schedule attached. The follow-up is almost always *and then what* - so have the three moves ready, and be clear that only the third ends the problem.
- The provider offers a larger size class that would fit today. Does that settle the question?Not on its own. It buys headroom, not a different ceiling - the tier still ends somewhere, and arriving on the largest class means the next refusal comes with no room left to manoeuvre. Ask what the largest class is and how long your measured growth takes to reach it. That interval is your planning horizon, and it is the number worth putting in the ticket.
- How do you keep serving while the question of which ceiling it is is being answered?Buy time deliberately rather than by accident: trim retention, move cold data out of the hot store, shed read traffic that does not need the primary, drop the least valuable load. Each buys months and none removes the ceiling. Write down which mitigations you spent, because headroom that was bought once cannot be bought again and the next conversation needs the truth.
saying these in an interview costs you the question
- Assumes any limit can be raised by opening a support case
- Treats the largest size class as proof that more capacity exists above it
- Plans a migration before confirming which kind of ceiling was hit
- Believes paying more always buys past a managed tier's design limit
- Reads a temporary throttling error as a permanent architectural cap