skip to content

questions

20

A scoring fleet takes 12,000 sensor readings per second at peak, each instance sustains 200, target utilisation 60% - how many instances?

level: juniorimportance: must knowfreq 62%

answer

  1. two divisions, not one
  2. peak rate, never the daily average
  3. saturation count comes first
  4. divide again by the utilisation target
  5. 12,000 / 200 / 0.6 = 100

basics

~10 s

One hundred instances. Dividing the 12,000-per-second peak by 200 gives 60 instances running flat out; dividing again by the 0.6 utilisation target gives 100, whose 20,000-per-second ceiling leaves peak sitting at 60%.

solid answer

~40 s

Two divisions, and they answer different questions. `12,000 / 200 = 60` is the *saturation* count: the fleet that exactly matches peak, which means every instance is busy 100% of the time and readings start queueing behind each other. The utilisation target converts that into a provisioned count, `60 / 0.6 = 100` instances, with a ceiling of `100 x 200 = 20,000` readings per second, so the 12,000 peak lands at 60% of it. Check it backwards before you hand the number in: the provisioned count is always larger than the saturation count, so if yours came out smaller you multiplied by the target instead of dividing. And size against the peak arrival rate, not the daily average - an average-rate fleet is under water for exactly the hours that matter.

go deeper

for a junior

Remember the shape: peak rate divided by per-instance throughput, then divided again by the utilisation target. Say the ceiling out loud as a check - 100 instances times 200 is 20,000 readings per second.

for a middle

Explain why the two divisions are different questions: the first gives a fleet that is 100% busy at peak, the second gives one with room for readings to wait briefly. Name where each input came from.

for a senior

Show that you distrust the inputs. State which window the peak was measured over, at what load the per-instance throughput was measured, and what the utilisation target is protecting - the scoring deadline, not the machines.

for a principal

Frame the utilisation target as the lever it is. Every point you move it trades machine count against the tail latency of a detection deadline, so it has to be owned by whoever owns that deadline, not chosen by the team that pays the bill.

## What the question is really asking Fleet sizing is a unit conversion. A rate the world imposes on you - sensor readings arriving per second at the scoring tier - and a rate one machine sustains - readings scored per second - become a count of machines. The arithmetic is two divisions, and an interviewer asking it is checking that you know what each division is for and which of the two numbers you refuse to hand in. Three inputs, and a sizing answer always needs all three: - **Peak arrival rate.** 12,000 readings per second: the highest rate the scoring tier is offered and expected to survive, not the average across the day. - **Sustained per-instance throughput.** 200 readings per second: what one instance holds indefinitely under production-shaped input, not a warm-cache burst. - **Utilisation target.** 0.6: the fraction of the provisioned fleet's ceiling that peak load is allowed to occupy. ## The first division gives a number you never provision `12,000 / 200 = 60`. Sixty instances have a combined ceiling of exactly 12,000 readings per second, so at peak they are busy 100% of the time. That is the **saturation count**: the floor below which the fleet cannot serve peak at all, not a fleet anyone runs. A tier at 100% utilisation has no idle slot for an arriving reading to land in, so readings wait behind one another and the wait keeps growing for as long as the peak lasts. ## The second division gives the provisioned count `60 / 0.6 = 100` instances. Dividing by the utilisation target is what converts a saturation count into a fleet you would actually run, and checking it backwards is how you catch the classic error: - fleet ceiling = `100 x 200 = 20,000` readings per second - utilisation at peak = `12,000 / 20,000 = 0.6`, which is the target The error is multiplying instead: `60 x 0.6 = 36` instances, a fleet whose ceiling is 7,200 readings per second, well under peak. The check is structural, so run it every time: **the provisioned count is always larger than the saturation count.** If your arithmetic produced a smaller one, you divided the wrong way round. | step | arithmetic | result | what it means | |---|---|---|---| | saturation count | `12,000 / 200` | 60 | 100% busy at peak, nothing spare | | provisioned count | `60 / 0.6` | 100 | the fleet you run | | fleet ceiling | `100 x 200` | 20,000/s | the rate it sustains | | utilisation at peak | `12,000 / 20,000` | 0.6 | the target, confirmed | ## Peak, not average Size against the peak arrival rate. The same sensor estate might average 4,000 readings per second across a day, which by the same arithmetic justifies `4,000 / 200 / 0.6 = 34` instances with a ceiling of 6,800 readings per second. That fleet is under water for every hour that matters: arrivals exceed the ceiling, the backlog grows for the length of the peak, and the anomaly the tier exists to catch is scored late. An average-rate figure answers a spend question. It never answers a capacity one. ## What the headroom actually buys The 40% the target leaves unoccupied is not waste and it is not a margin against arithmetic error. It is room for readings to wait briefly rather than at length: as a tier's utilisation climbs towards 1, the time a reading spends queued before it reaches a free worker slot grows out of all proportion to the load that caused it. Provisioning below saturation is how that wait is kept small enough for the scoring deadline to be met, which makes the utilisation target a latency decision written as a capacity number. ## Where each input goes wrong - **Peak read off an averaged chart.** A one-minute average hides the second-by-second peak the queue actually sees, and the queue is what decides latency. - **Per-instance throughput from a single-request benchmark.** Measured on an idle machine it is optimistic, and the fleet then runs hotter than the spreadsheet claims without anything alarming. - **A utilisation target copied from another service.** It has to be justified against this tier's own deadline and its own service-time variability. - **The model version left out of the record.** Per-instance throughput is a property of the artifact being served; promote a heavier one and the same 100 instances are a different fleet. Put the three inputs, the two divisions and the backwards check on the board in that order. The interviewer wants a number and the sentence that defends it.

  • The same estate averages 4,000 readings per second. How many instances does that justify, and why is it the wrong number to provision?
    Thirty-four: `4,000 / 200 / 0.6 = 33.3`, rounded up. But that fleet's ceiling is 6,800 readings per second and peak is 12,000, so for every peak hour arrivals exceed capacity, the backlog grows until the peak passes, and readings are scored minutes after the event they describe. Average sizing answers a spend question, never a capacity one.
  • If you raise the utilisation target from 60% to 80%, what does the fleet cost and what does it lose?
    `60 / 0.8 = 75` instances, a quarter fewer machines, with a ceiling of 15,000 readings per second. What it loses is queueing headroom: the wait factor `u / (1 - u)` moves from 1.5 to 4, so the average time a reading spends waiting for a free slot at peak roughly doubles and a half, and the p99 moves further still.
  • Peak is measured as 12,000 readings per second on a one-minute average. Why might the real sizing input be higher?
    Because a queue reacts to the instantaneous rate, not to a minute of it. A minute averaging 12,000 can contain ten-second stretches well above that, and during those stretches utilisation is above the target and a backlog forms. Take the peak from the shortest window that still reflects sustained load, typically seconds, and say which window you used.

saying these in an interview costs you the question

  • Sizes the fleet from the daily average arrival rate instead of the peak.
  • Multiplies the saturation count by the utilisation target instead of dividing by it.
  • Hands in 60 instances because that ceiling matches peak exactly.
  • Treats per-instance throughput as a fixed property of the machine class.
  • Says an autoscaler makes a provisioned count unnecessary.
open as a page

A bulk catalogue re-embedding job runs 400 accelerator-hours a month at $3 an hour for 20 million embeddings - what is the cost per embedding?

level: juniorimportance: must knowfreq 62%

basics

~10 s

Six hundredths of a cent. 400 hours at $3 is $1,200 of compute, divided by 20 million embeddings gives $0.00006 each, or $0.06 per thousand. That figure covers marginal compute only.

open as a page

In a video-moderation service, which half of the spend grows with upload volume: the quarterly training run or per-clip scoring?

level: juniorimportance: must knowfreq 62%

basics

~20 s

Per-clip scoring grows with upload volume; the quarterly training run does not. Training is paid once per model version, while scoring is paid again for every clip that arrives, so only the serving half is multiplied by traffic.

open as a page

A hosted invoice extractor charges per document — how do you find the monthly volume where running your own becomes cheaper?

level: middleimportance: must knowfreq 62%

basics

~20 s

Divide the fixed monthly spend of running your own by the per-document saving it earns: crossover volume = fixed monthly cost / (hosted price per document - your own marginal cost per document). Below that volume, buying is the cheaper of the two recurring bills.

open as a page

Which costs does a per-embedding figure computed from the catalogue re-embedding workers' compute bill alone leave out?

level: middleimportance: must knowfreq 58%

basics

~20 s

The amortised training run, the reserved-but-idle share of the accelerator fleet, vector storage and index writes, and orchestration and monitoring. On the worked catalogue refresh they turn $0.00006 of marginal compute into about $0.0006 fully loaded.

open as a page

A policy classifier costs 600 accelerator-hours per quarterly retrain and scores 1.9 billion clips at 40 ms each — which half dominates?

level: middleimportance: must knowfreq 57%

basics

~20 s

Serving dominates, by roughly 36 to 1. At 40 ms per clip, 1.9 billion clips need about 21,600 accelerator-hours against the retrain's 600, and the two halves would only be equal at about 54 million clips per retrain interval.

open as a page

Twenty senders carry most of your scanned-invoice volume — what recurring cost decides how long template rules stay the cheapest option?

level: seniorimportance: must knowfreq 56%

basics

~20 s

Human maintenance time. Per-sender template rules cost almost nothing to run, so the recurring bill is patch work: templates held multiplied by how often each sender changes its layout. It grows with senders covered, not with documents processed.

open as a page

Your anomaly-scoring fleet was sized so its 12,000-reading peak just fits the ceiling; nothing is dropped, yet p99 misses the 250 ms deadline - why?

level: seniorimportance: must knowfreq 57%

basics

~20 s

Queue wait. Sizing for throughput only checks that arrivals stay under the ceiling; a reading's latency is queue wait plus service time, and mean wait scales with utilisation as u / (1 - u), so a fleet at saturation queues heavily while its throughput still looks fine.

open as a page

In an anomaly-scoring fleet, where does one instance's 200-readings-per-second capacity come from, given 40 ms service time and eight worker slots?

level: middleimportance: should knowfreq 48%

basics

~20 s

Little's Law. Concurrency equals arrival rate times service time, so eight slots each held for 40 ms turn over 8 / 0.04 = 200 readings per second. The figure is arithmetic from a slot count and a service time, not a machine specification.

open as a page

A hosted invoice extractor prices per document — what else should the buy option be costed with?

level: seniorimportance: should knowfreq 47%

basics

~20 s

Everything the quoted rate hides: a throughput ceiling and a latency distribution you do not set, a model version that can change without any release of yours, and the integration, exception review and degraded path that stay yours whatever the price covers.

open as a page

A plant-wide fault makes every machine alarm at once and readings jump to 30,000 per second - what does autoscaling lag cost your 100-instance scoring fleet?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Roughly a five-minute backlog. A 20,000-per-second fleet offered 30,000 accumulates 10,000 readings a second, so a four-minute autoscaler lag builds 2.4 million queued readings that take minutes to drain - every one of them scored long after its operator window closed.

open as a page

Your per-embedding figure amortises a $90,000 training run over 18 months, but the embedding model is replaced after five - what breaks?

level: seniorimportance: should knowfreq 54%

basics

~10 s

The amortised share rises 3.6x, from $0.00025 to $0.0009 per embedding, carrying the fully-loaded figure from $0.0006 to $0.00125. Worse, the successor cannot reuse the old vectors, so the whole catalogue must be re-embedded.

open as a page

Skipping catalogue images whose bytes are unchanged cuts monthly re-embedding compute by 80% - what happens to cost per embedding?

level: seniorimportance: should knowfreq 44%

basics

~10 s

Marginal cost per embedding produced does not move: numerator and denominator fall together, so it stays $0.00006. Cost per listing kept current falls fivefold, and the fully-loaded figure per embedding actually rises.

open as a page

A policy change forces re-scanning a 4-billion-clip back catalogue — why is that spend neither the training half nor the per-upload half?

level: seniorimportance: should knowfreq 41%

basics

~20 s

A bulk re-scan is a third bucket: per-item like serving, but triggered by a policy decision rather than by traffic, and unbounded by any user-facing latency. At 40 ms per clip it costs about 44,400 accelerator-hours, roughly two quarters of live scoring in one pulse.

open as a page

A video-moderation scoring fleet is provisioned for peak; overnight uploads fall to a fifth of peak, so why does the monthly bill barely move?

level: seniorimportance: should knowfreq 52%

basics

~20 s

Because the serving half is paid as provisioned machine-hours, not as forward passes. Capacity held for the evening peak is billed through the quiet hours, so scoring fewer clips overnight consumes less compute without buying back any of the capacity already rented.

open as a page

Finance wants a three-year commitment to a hosted invoice extractor — which costs beyond the per-document rate decide whether to sign?

level: principalimportance: should knowfreq 40%

basics

~20 s

Exit cost, the asset you never accumulate, and where the documents are processed. A long commitment locks a rate you can compute and a coupling you cannot: downstream systems written against their fields, no labelled corpus of your own, and customer invoices leaving your boundary for three years.

open as a page

What makes buying hosted invoice extraction first and building later a real plan rather than a deferral?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

Three things: extractions and reviewer corrections stored in your own schema so a corpus accrues while you buy, a boundary that keeps exit cost flat, and a written trigger with a date. Without them, build-later is a wish and exit cost grows quietly.

open as a page

Your scoring tier's 40 ms per-reading service time was measured on an idle instance - what does that do to the fleet sized from it?

level: seniorimportance: nice to knowfreq 27%

basics

~20 s

It over-states capacity. Service time rises under concurrency, so if the real figure at eight busy slots is 55 ms the instance sustains 145 readings per second rather than 200, and the fleet built for 60% utilisation is actually running at 83%.

open as a page

How would you argue that $0.0006 per refreshed catalogue embedding is cheap, rather than just quoting the number?

level: seniorimportance: nice to knowfreq 31%

basics

~20 s

Put it against what a refreshed embedding earns, on the same denominator. Taking the refresh's attributed margin as $30,000 a month over 20 million embeddings, each costs $0.0006 and returns $0.0015 - and that return is very unevenly spread.

open as a page

A team proposes retraining a policy classifier yearly rather than quarterly to cut the accelerator bill — what caps that saving?

level: seniorimportance: nice to knowfreq 33%

basics

~20 s

The training half's share caps it. At roughly 2,400 of about 88,800 accelerator-hours a year, training is under 3% of compute, so dropping from four retrains to one saves at most about 2% — while the staleness it buys is paid on the serving side and outside the accelerator bill entirely.

open as a page