skip to content

Machine Profiles & Sizing

Machine families weighted toward processor, memory, local disk or accelerators, and sizing from measurement rather than from the server being replaced. Probed because rounding up is where bills grow.

on this pageshow

questions

4

You are moving a measured search service onto rented machines — which numbers set the machine size, and which do you ignore?

level: juniorimportance: must knowfreq 72%

answer

  1. measure the workload, not the box
  2. one full demand cycle of samples
  3. percentile, not average
  4. headroom for growth and a lost replica
  5. the old spec sheet is a purchase record

basics

~20 s

Size from the service's own observed processor, memory, disk and network use over a full demand cycle, taken at a high percentile with headroom added. The specification of the server being replaced records a purchase decision, not the workload, and is ignored.

solid answer

~50 s

Size from measurement of the workload, not from the box it is leaving. Collect processor time, resident memory, local disk throughput and network throughput over a full demand cycle — normally a week, so the weekly peak is inside it — and take a high percentile per dimension rather than an average, because the average hides the peak you must serve. Add explicit headroom: for growth before the next review, and for the moment one replica dies and the survivors absorb its share. Then pick the smallest offered size that clears every dimension. The old server's specification is evidence about a purchase made years ago, oversized on purpose for a depreciation period, and it records nothing about how much of itself was ever used. Re-measure once the service is running on the rented machine — its processor generation and storage path are not the ones you measured on.

code

pseudocode · 22 lines
pseudocode
# samples: one per minute, over a full demand cycle
peakCores      = percentile(samples.coreTime, 99)
peakMemoryGiB  = percentile(samples.residentMemoryGiB, 99)
peakNetOutMbps = percentile(samples.networkOutMbps, 99)

# survive losing one replica: survivors absorb its share
replicaFactor = replicaCount / (replicaCount - 1)

# room to grow before the next sizing review
growthFactor = 1.3

requiredCores  = peakCores * replicaFactor * growthFactor
requiredMemory = peakMemoryGiB * replicaFactor * growthFactor
requiredNetOut = peakNetOutMbps * replicaFactor * growthFactor

chosen = smallest size in offeredSizes where
           size.cores      >= requiredCores and
           size.memoryGiB  >= requiredMemory and
           size.netOutMbps >= requiredNetOut

# the dimension with the least slack is the one that forced the choice
binding = dimension of chosen with least slack

go deeper

for a junior

Recall that sizing starts from measurements of your own service — processor, memory, disk and network over a real week — and that the machine you are leaving tells you what somebody once bought, not what the service needs.

for a middle

Explain why a high percentile rather than a mean sets the size, and say what the headroom is actually for: growth before the next review, and absorbing a lost replica's share of the traffic.

for a senior

Show the loop you run in practice — measure, take the smallest size that clears every dimension, re-measure on the rented machine, correct once the demand cycle repeats — and name the dimension that bound the choice.

for a principal

Argue where the organisation's sizing standard sits: how much default headroom every team gets, who reviews an exception, and what a blanket habit of rounding up costs across a whole fleet.

## What sizing actually fixes Renting a machine means committing to a **profile** — a family weighting, and a size inside that family — before the workload has run there for a single minute. That one choice fixes two things at once: the quantity of processor time, memory, local disk throughput and network throughput the service may use, and the hourly charge you pay whether the service uses them or not. Sizing is the moment an engineering estimate turns into a recurring bill, which is why an interviewer asks how you arrived at the number rather than what the number was. ## Why the server you are replacing is not evidence The specification of an owned server is a record of a **purchase**, not a measurement of a **workload**. It was bought once, for a depreciation period measured in years, with margin for growth that may never have arrived and for a service that has since been rewritten. Its processor generation is old enough that the same work can need noticeably fewer cores on current hardware. And nothing on a specification sheet records how much of the machine was ever used: a box running at 8% average utilisation and one running at 80% look identical on paper. Copying that specification across imports somebody else's decision from years ago and then pays for it monthly. The honest starting point is what the service does, measured on the service. ## The measurement you collect Collect, per replica, over a **full demand cycle** — normally a week, because most services have a weekly shape, and longer where the business has a monthly or seasonal one: - **processor time**, as a fraction of a core, sampled finely enough that a one-minute spike is visible; - **resident memory**, the working set actually held, not the amount configured or reserved; - **local disk throughput and operation rate**, if the service reads or writes locally at all; - **network throughput in and out**, which on many rented sizes is itself capped by the size; - **the shape of the peak** — a brief spike, a daily plateau, or a monthly batch. Averages are the trap. A service averaging 20% of a core but holding 90% for the two hours that matter has to be sized for those two hours, because that is when customers are watching. Take a high percentile per dimension — the 95th or 99th of the samples — rather than the mean. ## Turning the measurement into a size 1. Take the high-percentile value for each dimension. 2. Multiply by a **replica factor**: if the work is spread over `N` replicas and you intend to survive losing one, the survivors each absorb an extra share, so scale by `N / (N - 1)` — for four replicas, about a third more each. 3. Multiply by a **growth factor** covering the period until you next expect to revisit the size. 4. Choose the smallest offered size that clears **every** dimension. The dimension that clears last is the binding one and picks the size; the others carry slack. 5. Re-measure on the rented machine and correct. | Input | Where it comes from | What it protects against | |---|---|---| | High percentile of use | a full demand cycle of samples | sizing for an average nobody experiences | | Replica factor | replica count and the failure you accept | one replica's loss saturating the rest | | Growth factor | the horizon until the next review | a resize forced by ordinary growth | | Binding dimension | whichever clears last | a size that is generous in the wrong resource | Step 5 is the one candidates skip. The rented machine has a different processor generation, a different storage path and a different network path from the machine you measured on, so identical work maps to a different utilisation. The pre-move numbers are an opening bid, not a result. ## What this size is not - **Not permanent.** A size is a rental. Treat the first one as a hypothesis with a review date attached. - **Not a latency promise.** A larger size adds quantity of resources; it does not necessarily make one unit of work finish faster. - **Not free insurance.** Every rung you climb is paid for every hour the machine exists. ## Where providers differ Platforms differ in how their size ladders step, in whether local disk arrives attached to the size or is rented separately, in whether network throughput scales with the size, and in whether changing a size requires the machine to be stopped first. Those differences change the mechanics of a resize, not the method: measure the workload, add explicit headroom, take the smallest size that clears it, and measure again once it is running.

  • Which percentile would you size from, and what decides it?
    High enough that the peak you promised to serve sits inside it, low enough that one pathological minute does not buy a permanently larger machine. The 95th is a common working choice for a service with a smooth peak; a spiky, latency-sensitive tier justifies the 99th. What decides it is what happens at the excluded moments: if exceeding the size degrades a customer-facing request, move up a percentile.
  • The service is brand new and has no measurements at all — now what?
    Get a number rather than guess one. Drive a load test at the demand you expect and measure that, or start deliberately small on metered capacity, observe a real cycle and resize. Say plainly that the first size is provisional, put a review date on it, and make sure the resize is cheap to perform before you need it.
  • Why re-measure on the rented machine instead of trusting the pre-move numbers?
    Because the rented machine is not the machine you measured on. A different processor generation, a different storage path and a different network path mean the same work lands at a different utilisation — sometimes lower, sometimes higher. Pre-move numbers pick the starting size; the post-move measurement is what you actually size from.

saying these in an interview costs you the question

  • Matches the replaced server's core and memory count one for one.
  • Sizes from average utilisation, leaving the real peak nowhere to go.
  • Samples one quiet afternoon and calls that a demand cycle.
  • Adds no headroom, so losing one replica saturates the survivors.
  • Treats the first size as final and never measures after the move.
open as a page

Machine families are weighted toward processor, memory, local disk or accelerators — which measurements pick one?

level: middleimportance: must knowfreq 58%

basics

~20 s

The ratio in a measured profile picks the family — which resource saturates first while the others idle. Memory-heavy work wants a memory-weighted family, compute-bound work a processor-weighted one, heavy local input and output a disk-weighted one, dense parallel numeric work an accelerator.

open as a page

An admin console on a fractional-core tier ran fine for weeks, then crawled during a campaign — why?

level: seniorimportance: should knowfreq 44%

basics

~20 s

A fractional-core machine accrues credit while it stays below its baseline share of a processor and spends that credit to burst above it. Weeks of idling built a balance; sustained campaign traffic drained it, and the machine dropped to its baseline share.

open as a page

Rented machine sizes usually step by doubling — what does rounding a service up to the next size buy?

level: middleimportance: nice to knowfreq 33%

basics

~20 s

Rounding up buys roughly double the resources for roughly double the hourly charge — a purchase, not a free safety margin. A service measured just past one rung spends most of the next rung idle, which is where a fleet's bill quietly grows.

open as a page