skip to content

Your nightly survey suite finishes fast on virtual targets and unpredictably on ruggedised handsets. What property of the fleet explains it?

level: seniorimportance: must knowfreq 56%

answer

  1. one capacity is bought, the other started
  2. an image duplicates; a unit does not
  3. supply shrinks as units die
  4. measure the wait per target
  5. hardware targets have a lifetime

basics

~20 s

Virtual capacity is manufactured on demand: an image started onto server capacity the provider extends by adding servers. A handset model exists as physical units nobody can duplicate, so that supply grows slowly, and it shrinks as units fail.

solid answer

~50 s

The two kinds of capacity are grown by completely different means. A virtual target is created from an image onto general-purpose server capacity, so the same image can be started again whenever compute is free, and the provider extends that capacity by adding servers. A physical handset cannot be created at all: the supply of a particular model on a particular platform version is however many units the provider bought, it shrinks as units fail or are retired, and for a ruggedised industrial handset that is no longer manufactured it cannot easily be grown at any price. So availability is a **per-target** property rather than a fleet-wide one, and neither your account's permissions nor your runner's worker count changes how many physical units exist. Measure availability per target, and ask the provider about the scarce ones rather than inferring it from a good week.

go deeper

for a junior

Know that a virtual target can be started whenever there is server capacity, while a real device has to be free, and that rare models are the ones that keep you waiting.

for a middle

Be ready to explain why physical supply cannot be grown quickly, and why a model paired with a platform version is a narrower target than the model on its own.

for a senior

Expect to diagnose an unpredictable hardware lane: what you would measure per target, over what window, and how you would separate a supply constraint from an account limit or a broken harness.

for a principal

Be ready to treat scarce hardware as a dependency with a lifetime: which coverage rests on it, what you ask the provider and record, and what the plan is when the model leaves the fleet.

A hosted fleet presents both kinds of target through the same remote endpoint, so from your suite's point of view asking for one looks exactly like asking for the other. The behaviour diverges because the two kinds of capacity are produced by completely different means, and the divergence shows up first as unpredictable availability on the hardware lane. ## Where each kind of capacity comes from A virtual target is **manufactured on demand**. The provider holds an image and general-purpose server capacity, and creating an instance is a scheduling decision: find free compute, start the image. The same image can back many instances at once, and the provider grows the underlying capacity the way anyone grows server capacity, by adding servers. A physical target is **bought**. Someone acquired that model, racked it, wired it for power and control, and that unit now exists. The provider can grow this supply only by buying and racking more of the same model, which is slower than starting a process, and for some models is not possible at all. ## Why a physical model's supply is hard to grow, and easy to lose - The model may no longer be manufactured, so additional units come only from the second-hand market, if anywhere. - A ruggedised industrial handset is a niche product with a small installed base to begin with. - Racking a unit is physical work with a lead time, not a configuration change. - Units leave: batteries swell, screens fail, ports wear out, and a dead unit is removed from the pool. - A given model on a given platform version is a narrower target than the model alone, and a unit updated in place stops satisfying the older version. The last point is worth dwelling on, because it is the one people miss. Your matrix cell is usually a pairing of model and platform version. The population satisfying that pairing can shrink without any unit failing, simply because units moved to a newer platform version. ## Availability is a per-target property The mistake this scenario exposes is treating availability as one number for the fleet. It is not. Your virtual lane and your hardware lane have structurally different availability, and within the hardware lane a common current model and a rare ruggedised one behave differently again. That has practical consequences for how you operate the suite: - Record time-to-session **per target**, not as one average across the run, because the average is dominated by the plentiful targets and hides the scarce one. - Record it over a long enough window to see the bad weeks, since a scarce target can be free for weeks and then not. - Treat a slow start on a scarce target as a supply observation rather than as an incident, and a slow start on the virtual lane as a genuine anomaly worth investigating. - Know which cell of your matrix is the scarce one before somebody asks why the nightly run has not finished. There is a further limit worth naming precisely, because it is easy to conflate. How many sessions your account may hold at once, and how many workers your runner starts, are both real constraints, and they are different questions from this one. Physical unit count is a third limit underneath both: even where your account permits more and your runner has more workers ready, the hardware has to exist and be free. ## What you can measure, and what you must ask Measurable from your own side: 1. Time from requesting a target to holding a usable session, tagged with the exact target asked for. 2. How often a request for a scarce target is satisfied at all within your run's window. 3. Which targets your suite actually asks for, which is frequently not the list anybody believes it is. Answerable only by the provider, so ask and record it: - Whether a model you depend on is still being acquired, or is being allowed to dwindle. - Whether units of that model may be updated in place to a newer platform version. - Whether the model is scheduled to leave the fleet, and with what notice. Nothing in that second list is inferable from your own telemetry, and guessing at it is exactly the habit that turns a supply constraint into a surprise. What you can establish, you establish; what you cannot, you ask about, and you record the answer with a date on it. ## Planning for a target that leaves An image can in principle be kept and started again long after the software it contains is obsolete. A handset cannot: it physically wears out and eventually leaves the fleet, whether or not anyone still wants to test on it. So a hardware lane has an end-of-life problem that a virtual lane largely does not. Sound practice is to treat the scarce hardware target as a **dependency with a lifetime**: know which cell depends on it, know what you would do if it became unavailable next quarter, and make sure that answer is not discovered on the night it happens. That is a supply question about the fleet you rent, and it is worth asking long before it becomes an availability incident.

  • Why can a model and platform version pairing get scarcer without any unit failing?
    Because the pairing, not the model, is what your matrix cell names. If units in the fleet are updated in place to a newer platform version, they stop satisfying a request that pins the older one, and the population matching your cell shrinks while the physical count is unchanged. That is why it is worth asking whether units may be updated in place, and why a suite should assert the platform version it actually received.
  • What would you instrument to see this coming rather than discovering it at midnight?
    Time-to-session tagged with the exact target requested, held over a long window so weekly and seasonal variation is visible, plus the share of requests for each scarce target that were satisfied inside the run's window. Trend those per target rather than in aggregate. A scarce target usually degrades gradually as units leave the pool, so the trend is visible well before it becomes a failed run.
  • How is this different from your account simply being allowed fewer sessions at once?
    They are separate limits that happen to produce the same symptom. An account-level ceiling is a permission the provider sets and can change, and it applies across targets. Physical unit count is a fact about the world: it constrains one specific target and no permission raises it. Diagnosing the wrong one wastes time, so confirm which limit you are hitting before proposing a remedy.

saying these in an interview costs you the question

  • Treats availability as one number for the whole fleet
  • Assumes more workers will speed up a scarce hardware lane
  • Believes a provider can add units of any model on request
  • Confuses an account-level limit with the number of physical units
  • Never asks whether a depended-on model is leaving the fleet