skip to content

In an anomaly-scoring fleet, where does one instance's 200-readings-per-second capacity come from, given 40 ms service time and eight worker slots?

level: middleimportance: should knowfreq 48%

answer

  1. concurrency is the missing link
  2. in flight, not per second
  3. arrival rate times service time
  4. slots divided by service time
  5. 8 / 0.04 = 200 per second

basics

~20 s

Little's Law. Concurrency equals arrival rate times service time, so eight slots each held for 40 ms turn over 8 / 0.04 = 200 readings per second. The figure is arithmetic from a slot count and a service time, not a machine specification.

solid answer

~50 s

Little's Law says that for any stable system the items inside it equal the arrival rate times the time each one spends inside: `L = lambda x W`. Read it one way and it sizes an instance: eight slots held for 40 ms of work each give `8 / 0.04 = 200` readings per second. Read it the other way and it sizes the fleet: a 12,000-per-second peak at 40 ms of service time needs `12,000 x 0.04 = 480` readings in service at any instant, which is `480 / 8 = 60` saturated instances and 100 at a 0.6 utilisation target - the same answer from the other direction, which is why it is worth saying both out loud. The `W` in that formula is **service time**, the work on one reading with no waiting in it. Substitute an observed end-to-end latency and the arithmetic silently over-provisions.

code

pseudocode · 17 lines
pseudocode
peakArrivalRate   = 12000     // readings per second at peak
serviceTime       = 0.040     // seconds of work per reading, no queue wait
slotsPerInstance  = 8
utilisationTarget = 0.60

// Little's Law: readings in service = arrival rate x service time
concurrency = peakArrivalRate * serviceTime               // 480 readings, a count

saturationInstances = ceil(concurrency / slotsPerInstance)        // 60
provisionedInstances = ceil(saturationInstances / utilisationTarget)  // 100

// check it backwards before reporting the number
fleetCeiling      = provisionedInstances * slotsPerInstance / serviceTime  // 20000 per second
utilisationAtPeak = peakArrivalRate / fleetCeiling                        // 0.60

assert provisionedInstances > saturationInstances
assert utilisationAtPeak <= utilisationTarget

go deeper

for a junior

Learn the two forms: readings in flight equals rate times service time, and a slot count divided by service time gives throughput. Keep the units straight - concurrency is a count, throughput is a rate.

for a middle

Be able to run it in both directions and get the same fleet, and be able to say why service time is the only one of the three durations on a serving path that belongs in the formula.

for a senior

Show where the service time came from. A sizing built on a number measured at the wrong load or on a previous model version is arithmetic that is right about the wrong tier.

for a principal

Treat concurrency as the contract between teams. It is the one number that ties an arrival-rate forecast, a model artifact's cost and a machine count together, so it is what a capacity review should argue about.

## Little's Law, stated **Little's Law** says that for any stable system the average number of items inside it equals the average arrival rate multiplied by the average time each item spends inside: `L = lambda x W`. It is distribution-free. It assumes nothing about how arrivals are spaced or how long service takes - only that the system is stable, meaning nothing accumulates without bound, and that you average over a long enough window. That generality is exactly why it is the first tool anyone reaches for in a sizing conversation: you can apply it without knowing anything about the traffic's shape. On a model-serving tier you apply it twice, in two directions. ## Direction one: what one instance sustains An instance has eight worker slots. A reading occupies one slot for a **service time** of 40 ms - the feature assembly and the model's forward pass, the work itself, with no queue wait in it. Rearranged for the service stage, `lambda = L / W`: - `L` is 8, the readings in service when every slot is busy - `W` is 0.04 seconds, the service time - `lambda = 8 / 0.04 = 200` readings per second That 200 is not a specification; it is the arithmetic consequence of a slot count and a service time, and it moves when either does. Halve the service time and the same instance sustains 400 per second. Drop to four slots and it sustains 100. ## Direction two: what the fleet must hold Now run it forwards. At a peak of 12,000 readings per second with a 40 ms service time, the concurrency the fleet must be able to hold is `L = 12,000 x 0.04 = 480` readings in service at any instant. Divide by the eight slots an instance offers and you are back at 60 saturated instances, then 100 at a 0.6 utilisation target. Two routes, one answer - and getting a different one means one of the inputs is being used inconsistently. | quantity | symbol | this tier | what it is | |---|---|---|---| | peak arrival rate | `lambda` | 12,000/s | readings offered per second | | service time | `W` | 40 ms | work on one reading, no waiting | | concurrency in service | `L` | 480 | readings being scored at an instant | | slots per instance | - | 8 | one instance's concurrency | | per-instance throughput | - | 200/s | `8 / 0.04` | Notice what `L` is and is not. It is a **count** - 480 readings, dimensionless - not a rate. Candidates who report "480 readings per second" have lost the units, and the error usually propagates into a fleet off by a factor of the service time. ## Service time is not latency The most common way to get this wrong is to substitute the wrong `W`. Three durations on a serving path all get called latency and only one belongs in this arithmetic: | duration | what it includes | belongs in the sizing arithmetic? | |---|---|---| | service time | the work done on one reading | yes - this is `W` | | queue wait | time spent waiting for a free slot | no | | end-to-end latency | queue wait plus service time, plus transport | no | Size from an observed end-to-end p99 and you provision a fleet several times larger than the work requires, then watch it sit at a utilisation nothing predicted. Service time has to come from a timer that starts when the reading takes a slot and stops when it releases it. ## What Little's Law does not tell you It gives you concurrency, and concurrency gives you a machine count. It says nothing at all about **how long a reading waits** before it gets a slot. That is governed by how bursty arrivals are and how variable service times are, which is the territory of queueing approximations such as **Kingman's formula**, and it is why sizing does not stop at the saturation count. It also assumes stability. Apply it to a tier whose arrival rate exceeds its ceiling and the `L` it reports grows without bound: that is the arithmetic telling you a backlog is accumulating, not telling you to provision an ever-larger number of slots. ## Using it in the room 1. State the inputs you were given and name the one you are assuming: arrival rate, service time, slots per instance. 2. Compute `lambda x W` and say out loud that the result is a count of readings in flight, not a rate. 3. Divide by slots per instance for the saturation count, then by the utilisation target for the provisioned count. 4. Check backwards - provisioned count times per-instance throughput should give a ceiling that leaves peak at the target. 5. Say where the service time came from and at what load it was measured. That sentence is what separates a sizing answer from a division.

  • Does Little's Law assume anything about how the readings arrive?
    No. It holds for any stable system averaged over a long enough window, whatever the arrival or service-time distributions - that is what makes it safe to use with no traffic model at all. The distributions do matter, but to a different quantity: they set how long a reading waits for a free slot, which Little's Law never claims to tell you.
  • The fleet holds 480 readings in service at peak. Where are the rest of the slots?
    Idle, by design. One hundred instances offer 800 slots and peak occupies 480 of them, which is the 60% utilisation target expressed as a count. Apply Little's Law to the whole tier rather than just the service stage and the number grows: it becomes the arrival rate times queue wait plus service time, so the readings waiting for a slot are counted too.
  • An instance is given sixteen slots instead of eight, on the same hardware. Does its throughput double?
    Only if the machine can genuinely execute sixteen readings at once. Slots above the available parallelism do not add throughput; they move the queue from outside the process to inside it, and per-reading service time rises as the extra work contends for the same cores and memory. Little's Law still holds - it is the 40 ms that stops being 40 ms.

saying these in an interview costs you the question

  • Reports the concurrency figure as a rate per second rather than a count.
  • Uses the observed end-to-end p99 as the service time in the arithmetic.
  • Claims Little's Law needs Poisson arrivals to be valid.
  • Thinks doubling the slots per instance halves the per-reading service time.
  • Expects Little's Law to predict how long a reading waits for a slot.