A scoring fleet takes 12,000 sensor readings per second at peak, each instance sustains 200, target utilisation 60% - how many instances?
answer
- two divisions, not one
- peak rate, never the daily average
- saturation count comes first
- divide again by the utilisation target
- 12,000 / 200 / 0.6 = 100
basics
~10 sOne hundred instances. Dividing the 12,000-per-second peak by 200 gives 60 instances running flat out; dividing again by the 0.6 utilisation target gives 100, whose 20,000-per-second ceiling leaves peak sitting at 60%.
solid answer
~40 sTwo divisions, and they answer different questions. `12,000 / 200 = 60` is the *saturation* count: the fleet that exactly matches peak, which means every instance is busy 100% of the time and readings start queueing behind each other. The utilisation target converts that into a provisioned count, `60 / 0.6 = 100` instances, with a ceiling of `100 x 200 = 20,000` readings per second, so the 12,000 peak lands at 60% of it. Check it backwards before you hand the number in: the provisioned count is always larger than the saturation count, so if yours came out smaller you multiplied by the target instead of dividing. And size against the peak arrival rate, not the daily average - an average-rate fleet is under water for exactly the hours that matter.
go deeper
Remember the shape: peak rate divided by per-instance throughput, then divided again by the utilisation target. Say the ceiling out loud as a check - 100 instances times 200 is 20,000 readings per second.
Explain why the two divisions are different questions: the first gives a fleet that is 100% busy at peak, the second gives one with room for readings to wait briefly. Name where each input came from.
Show that you distrust the inputs. State which window the peak was measured over, at what load the per-instance throughput was measured, and what the utilisation target is protecting - the scoring deadline, not the machines.
Frame the utilisation target as the lever it is. Every point you move it trades machine count against the tail latency of a detection deadline, so it has to be owned by whoever owns that deadline, not chosen by the team that pays the bill.
## What the question is really asking Fleet sizing is a unit conversion. A rate the world imposes on you - sensor readings arriving per second at the scoring tier - and a rate one machine sustains - readings scored per second - become a count of machines. The arithmetic is two divisions, and an interviewer asking it is checking that you know what each division is for and which of the two numbers you refuse to hand in. Three inputs, and a sizing answer always needs all three: - **Peak arrival rate.** 12,000 readings per second: the highest rate the scoring tier is offered and expected to survive, not the average across the day. - **Sustained per-instance throughput.** 200 readings per second: what one instance holds indefinitely under production-shaped input, not a warm-cache burst. - **Utilisation target.** 0.6: the fraction of the provisioned fleet's ceiling that peak load is allowed to occupy. ## The first division gives a number you never provision `12,000 / 200 = 60`. Sixty instances have a combined ceiling of exactly 12,000 readings per second, so at peak they are busy 100% of the time. That is the **saturation count**: the floor below which the fleet cannot serve peak at all, not a fleet anyone runs. A tier at 100% utilisation has no idle slot for an arriving reading to land in, so readings wait behind one another and the wait keeps growing for as long as the peak lasts. ## The second division gives the provisioned count `60 / 0.6 = 100` instances. Dividing by the utilisation target is what converts a saturation count into a fleet you would actually run, and checking it backwards is how you catch the classic error: - fleet ceiling = `100 x 200 = 20,000` readings per second - utilisation at peak = `12,000 / 20,000 = 0.6`, which is the target The error is multiplying instead: `60 x 0.6 = 36` instances, a fleet whose ceiling is 7,200 readings per second, well under peak. The check is structural, so run it every time: **the provisioned count is always larger than the saturation count.** If your arithmetic produced a smaller one, you divided the wrong way round. | step | arithmetic | result | what it means | |---|---|---|---| | saturation count | `12,000 / 200` | 60 | 100% busy at peak, nothing spare | | provisioned count | `60 / 0.6` | 100 | the fleet you run | | fleet ceiling | `100 x 200` | 20,000/s | the rate it sustains | | utilisation at peak | `12,000 / 20,000` | 0.6 | the target, confirmed | ## Peak, not average Size against the peak arrival rate. The same sensor estate might average 4,000 readings per second across a day, which by the same arithmetic justifies `4,000 / 200 / 0.6 = 34` instances with a ceiling of 6,800 readings per second. That fleet is under water for every hour that matters: arrivals exceed the ceiling, the backlog grows for the length of the peak, and the anomaly the tier exists to catch is scored late. An average-rate figure answers a spend question. It never answers a capacity one. ## What the headroom actually buys The 40% the target leaves unoccupied is not waste and it is not a margin against arithmetic error. It is room for readings to wait briefly rather than at length: as a tier's utilisation climbs towards 1, the time a reading spends queued before it reaches a free worker slot grows out of all proportion to the load that caused it. Provisioning below saturation is how that wait is kept small enough for the scoring deadline to be met, which makes the utilisation target a latency decision written as a capacity number. ## Where each input goes wrong - **Peak read off an averaged chart.** A one-minute average hides the second-by-second peak the queue actually sees, and the queue is what decides latency. - **Per-instance throughput from a single-request benchmark.** Measured on an idle machine it is optimistic, and the fleet then runs hotter than the spreadsheet claims without anything alarming. - **A utilisation target copied from another service.** It has to be justified against this tier's own deadline and its own service-time variability. - **The model version left out of the record.** Per-instance throughput is a property of the artifact being served; promote a heavier one and the same 100 instances are a different fleet. Put the three inputs, the two divisions and the backwards check on the board in that order. The interviewer wants a number and the sentence that defends it.
- The same estate averages 4,000 readings per second. How many instances does that justify, and why is it the wrong number to provision?Thirty-four: `4,000 / 200 / 0.6 = 33.3`, rounded up. But that fleet's ceiling is 6,800 readings per second and peak is 12,000, so for every peak hour arrivals exceed capacity, the backlog grows until the peak passes, and readings are scored minutes after the event they describe. Average sizing answers a spend question, never a capacity one.
- If you raise the utilisation target from 60% to 80%, what does the fleet cost and what does it lose?`60 / 0.8 = 75` instances, a quarter fewer machines, with a ceiling of 15,000 readings per second. What it loses is queueing headroom: the wait factor `u / (1 - u)` moves from 1.5 to 4, so the average time a reading spends waiting for a free slot at peak roughly doubles and a half, and the p99 moves further still.
- Peak is measured as 12,000 readings per second on a one-minute average. Why might the real sizing input be higher?Because a queue reacts to the instantaneous rate, not to a minute of it. A minute averaging 12,000 can contain ten-second stretches well above that, and during those stretches utilisation is above the target and a backlog forms. Take the peak from the shortest window that still reflects sustained load, typically seconds, and say which window you used.
saying these in an interview costs you the question
- Sizes the fleet from the daily average arrival rate instead of the peak.
- Multiplies the saturation count by the utilisation target instead of dividing by it.
- Hands in 60 instances because that ceiling matches peak exactly.
- Treats per-instance throughput as a fixed property of the machine class.
- Says an autoscaler makes a provisioned count unnecessary.