Machine families are weighted toward processor, memory, local disk or accelerators — which measurements pick one?
answer
- the ratio, not the absolute number
- which resource saturates first
- family fixes the ratio, size scales it
- balanced is a default, not an answer
- accelerators only for dense parallel numeric work
basics
~20 sThe ratio in a measured profile picks the family — which resource saturates first while the others idle. Memory-heavy work wants a memory-weighted family, compute-bound work a processor-weighted one, heavy local input and output a disk-weighted one, dense parallel numeric work an accelerator.
solid answer
~50 sA family is a **ratio** of resources; a size is a quantity of that ratio. So the family is picked by which resource saturates first in the measured profile while the others sit idle. Processor near its limit with memory barely touched points at a processor-weighted family; memory near full with idle cores points at a memory-weighted one; a workload whose limit is local disk throughput or operation rate points at a family with fast attached local disk; dense parallel numeric work points at an accelerator-weighted family. A balanced family is the right answer only when the measurements really are balanced, not a default for everything. The consequence worth stating out loud is that you cannot fix a ratio problem by taking a larger size: a size step scales both resources together, so you would buy the cores you do not need in order to reach the memory you do.
go deeper
Remember that machine families differ in the ratio between processor, memory, disk and accelerator, and that you choose among them by looking at which resource your own service runs out of first.
Explain that a family sets the ratio while a size sets the quantity, and show why climbing the size ladder cannot cure a ratio mismatch — it multiplies the resource you do not need alongside the one you do.
Demonstrate reading percentiles per dimension to find the binding resource, spotting the case where one machine is really carrying two different profiles, and re-checking the family after a release that changes the service's shape.
Take a position on how many families a fleet should standardise on: fewer families simplify purchasing and capacity planning, more families fit workloads tightly. Say where you draw that line and what evidence would move it.
## Family is a ratio, size is a quantity Every provider sells rented machines as a small number of **families**, each with a characteristic ratio between processor, memory, local disk and any attached accelerator, plus a ladder of **sizes** inside each family that scales that ratio up and down. Choosing a profile is therefore two decisions, and they answer different questions: - the **family** answers *what shape does this workload have?* - the **size** answers *how much of that shape does it need?* Getting the second wrong costs money on one axis. Getting the first wrong costs money on every axis at once, because the only way to buy more of the resource you are short of is to buy more of everything. ## Reading the profile The input is the same measurement used for sizing — processor time, resident memory, local disk throughput and operation rate, network throughput — but read as a **ratio** rather than as absolute numbers. The question is which dimension reaches its limit first while the others are still slack. | Family weighting | The measured signature | Workloads that usually land there | |---|---|---| | Balanced | processor and memory rise together | request-serving tiers, small services, most first deployments | | Processor-weighted | processor saturates while memory stays low | encoding, compression, parsing, simulation, hot request paths | | Memory-weighted | memory near full while cores idle | in-memory indexes and caches, large joins and aggregations | | Local-disk-weighted | local disk throughput or operation rate is the ceiling | log and event processing, local scratch space, shuffle-heavy jobs | | Accelerator-weighted | dense parallel numeric work the processors cannot keep up with | model training and inference, some numeric and media pipelines | A customer-facing search tier is the interesting case, because it can land in two different rows depending on how it is built: an index held entirely in memory makes it memory-weighted, while an index read from local disk on every query makes it disk-weighted. The measurement decides, not the label on the service. ## Why a bigger size does not fix a wrong family This is the mechanical point underneath the question. Inside a family, the ladder multiplies **both** resources together. A service needing far more memory than processor, sized inside a balanced family, buys processor it will never use in order to reach the memory it needs — and pays for the whole machine every hour it exists. The memory-weighted family exists precisely to sell that ratio directly. The same argument runs the other way: a compute-bound job parked on a memory-weighted family rents memory it never fills. ## Choosing when nothing fits cleanly Real profiles rarely match a family exactly. In order of preference: 1. Pick the family whose scarce resource is the one binding you, and accept measurable slack on the others. 2. Check whether the profile is really two workloads sharing one machine — a memory-hungry cache beside a processor-hungry request path — in which case splitting them lets each take the family it wants. 3. Only then pay for slack deliberately, and record why, so the next review does not re-litigate it from scratch. ## Accelerators are a different decision An accelerator-weighted family is the one where the most expensive component sits idle unless the software is written to use it. The test is not *is this workload important* or *is this workload slow*, but *is the work dense, parallel and numeric, and does the code actually dispatch it to the accelerator?* Work that is slow because it waits on the network, or because it is one long dependent chain of operations, gains nothing from an accelerator while paying the most per hour of any family for the privilege. ## Practical cautions - **Read percentiles per dimension, not one summary number.** A short disk-bound phase inside an otherwise processor-bound job disappears entirely from an average. - **Watch for a limit you did not measure.** Network throughput and local disk operation rate are frequently capped by the size rather than by the hardware, so a profile can be bound by an allowance rather than by a resource. - **Re-read the ratio after any significant change to the service.** Adding an in-memory cache can move a service from one family to another in a single release. ## Where providers differ Platforms differ in how many families they offer, in whether local disk is part of the family or rented separately, in whether an accelerator can be attached to an otherwise ordinary family, and in how finely the ladder steps. What does not differ is the method: measure the ratio, name the resource that binds, and choose the family that sells that ratio.
- A service needs far more memory than processor. Why not simply rent a larger size in a balanced family?Because a size step inside a family scales both resources together. To reach the memory you need you buy processor you will never use, and pay for the whole machine every hour it exists. A memory-weighted family sells that ratio directly, so the same memory arrives without the surplus cores attached to it.
- The profile does not match any family cleanly. How do you choose?Choose the family whose scarce resource is the one binding you, and accept measurable slack on the rest. Before paying for that slack permanently, check whether the profile is really two workloads on one machine — a memory-hungry cache beside a processor-hungry request path — because splitting them often lets each take the family it actually wants.
- What would make you reject an accelerator-weighted family for work that is genuinely slow?Slowness that is not dense parallel numeric work. If the service is waiting on the network, or running one long dependent chain of operations, the accelerator sits unused while being the most expensive component on the machine. The other requirement is that the code actually dispatches work to it; an accelerator nothing addresses is pure cost.
saying these in an interview costs you the question
- Picks a family from the workload's importance rather than its measured ratio.
- Believes a larger size can fix a wrong ratio; it scales both resources.
- Uses a balanced family for everything and buys memory by buying cores.
- Asks for accelerators because the work is slow, not because it is parallel numeric.
- Reads only averages, so a short disk-bound phase never appears at all.