skip to content

Your scoring tier's 40 ms per-reading service time was measured on an idle instance - what does that do to the fleet sized from it?

level: seniorimportance: nice to knowfreq 27%

answer

  1. measured idle, run busy
  2. contention is charged to service time
  3. the ceiling was never what you thought
  4. utilisation 83%, not 60%
  5. service time belongs to the model version

basics

~20 s

It over-states capacity. Service time rises under concurrency, so if the real figure at eight busy slots is 55 ms the instance sustains 145 readings per second rather than 200, and the fleet built for 60% utilisation is actually running at 83%.

solid answer

~50 s

Service time is not a constant of the machine; it is what one reading costs *at the concurrency you intend to run*. Eight readings sharing an instance contend for cores, memory bandwidth, caches and the connection pool used to fetch features, so the per-reading figure measured one request at a time is always optimistic. Suppose the honest number at eight busy slots is 55 ms. Then per-instance throughput is `8 / 0.055 = 145` readings per second, the 100-instance fleet's ceiling is 14,500 rather than 20,000, and at a 12,000 peak true utilisation is `12,000 / 14,500 = 83%`, not the 60% the plan claimed. Nothing alarms, because throughput still fits - the headroom simply is not there, and the p99 sits on the steep part of the queueing curve. Measure service time on a loaded instance, with production-shaped readings, and re-measure it for every model version.

go deeper

for a junior

Remember that a timing taken on an idle machine is not what you get when eight readings run at once. Always ask at what load a performance number was measured.

for a middle

Explain what contends - cores, memory bandwidth, the feature-fetch client, allocation - and redo the arithmetic with the honest figure to show how the fleet's real ceiling and utilisation move.

for a senior

Point out that none of the usual alarms see this: throughput still fits and nothing is dropped, so it surfaces as an unexplained latency regression. Ask for the latency-against-load curve, not a single number.

for a principal

Make service time part of the promotion contract. When the sizing input is a property of the artifact, a model release changes the fleet's capacity, and that has to be reviewed with the quality numbers rather than discovered afterwards.

## The number you measured is not the number you need The sizing arithmetic takes a per-reading service time and turns it into a machine count. That makes service time the most load-bearing input in the calculation, and the easiest one to measure badly: fire one reading at an otherwise idle instance, time it, and you have a number that is precise, repeatable and wrong for the fleet you are about to build. What the arithmetic needs is the service time **at the concurrency you intend to run**, because that is the regime the fleet will live in. ## What gets slower when the instance is busy Eight readings in flight on one instance are not eight independent machines. They contend, and the contention shows up as per-reading time: - **Compute.** Worker slots share a fixed number of cores; above the real parallelism, slots wait rather than run. - **Memory bandwidth and caches.** A model's working set evicted by a neighbouring reading has to be re-fetched, and that cost lands inside service time. - **The feature fetch.** Slots share one client and one connection pool to whatever serves the features; at concurrency, queuing inside that client is charged to each reading's service time. - **Allocation and housekeeping.** Memory management, metrics emission and logging all scale with request rate and steal time from the work. - **The input mix.** A benchmark usually replays one shape of reading. Production has a distribution, and the heavy tail of it is served at the same concurrency as the rest. None of these is exotic. The point is that they are all invisible to a single-request measurement. ## Recomputing the fleet honestly | quantity | from the idle measurement | from the loaded measurement | |---|---|---| | service time | 40 ms | 55 ms | | per-instance throughput | `8 / 0.040` = 200/s | `8 / 0.055` = 145/s | | ceiling of 100 instances | 20,000/s | 14,500/s | | utilisation at a 12,000 peak | 60% | 83% | | queue-wait factor `u / (1 - u)` | 1.5 | ~4.9 | The fleet did not change. The plan's belief about it did. A tier the design says has 40% of its ceiling spare is in fact running with 17%, and the wait factor that drives the p99 has more than tripled - which is why this shows up first as a latency complaint and never as a capacity alert. ## Service time belongs to the model version There is a second reason a measured service time goes stale, and it is specific to a model-serving tier: **the number is a property of the artifact being served, not of the fleet.** Promote a version with more parameters, more feature lookups or a heavier pre-processing step and per-reading service time moves, which silently re-sizes the fleet without anybody touching the fleet. A promotion that raises service time from 40 ms to 52 ms cuts per-instance throughput to `8 / 0.052 = 154` readings per second and the fleet's ceiling to about 15,400. Peak still fits, so no throughput alarm fires; utilisation quietly moves from 60% to 78% and the p99 with it. Treat measured service time as part of what a candidate model must report before it is promoted, alongside its quality numbers. ## Measuring it honestly 1. Replay production-shaped readings, in their real mix of sizes, not a single synthetic shape. 2. Drive the instance to the concurrency you intend to run at and hold it there long enough for caches and pools to reach steady state. 3. Take service time from a timer that starts when the reading occupies a worker slot, so queue wait is excluded. 4. Validate by pushing to saturation: measured maximum throughput should match slots divided by the service time you recorded. A gap means the timer is measuring the wrong span. 5. Record the number with the model version and the hardware class it was measured on, and repeat it for every promotion. ## Why nothing alarms This failure hides because every signal that would catch it is watching the wrong quantity. Throughput fits, so no capacity alarm fires. Nothing is dropped, so no error alarm fires. Utilisation may well be reported against a stale assumed ceiling, so even the utilisation panel agrees with the plan. The only signals that see it are a measured p99 and the instance's own saturation behaviour - which is the argument for treating the latency-against-load curve, rather than a single service-time number, as the artifact that sizing is built on.

  • A promoted model version raises per-reading service time from 40 ms to 52 ms. What happens to the fleet you already provisioned?
    Per-instance throughput falls to `8 / 0.052 = 154` readings per second and the 100-instance ceiling to about 15,400. Peak of 12,000 still fits, so nothing is dropped and no throughput alarm fires - but utilisation at peak moves from 60% to roughly 78%, and the queue-wait factor from 1.5 to about 3.5, which shows up as a p99 regression nobody attributes to the promotion.
  • How do you measure service time so the sizing arithmetic is honest?
    Drive a representative instance to the concurrency you intend to run at, with production-shaped readings, and hold it in steady state. Time from the moment a reading takes a worker slot, so queue wait is excluded. Then validate by pushing to saturation and checking that measured maximum throughput equals slots divided by the recorded service time.

saying these in an interview costs you the question

  • Benchmarks one reading at a time and calls the result service time.
  • Treats per-reading service time as a property of the hardware alone.
  • Uses observed end-to-end latency as service time, double-counting queue wait.
  • Trusts the utilisation figure the sizing plan predicted rather than a measured one.
  • Expects a throughput alarm to catch a service-time regression.