skip to content

Should a test estate pin seeds, timezone and locale everywhere, or vary them in some runs?

level: principalimportance: should knowfreq 36%

answer

  1. Two runs with two different jobs
  2. Red must mean something in the gate
  3. Production is the varying environment
  4. One dimension at a time
  5. Fund the triage before enabling it

basics

~20 s

Pin them in the blocking run, where a failure must reproduce before it blocks anyone, and vary them only in a separate scheduled run that prints what it chose. Total pinning finds no latent assumptions; unmanaged variation reads as flakiness.

solid answer

~50 s

Treat it as two tiers with different jobs. The blocking pre-merge run is fully pinned: a failure there stops other people's work, so it must reproduce on demand and must never be caused by which machine picked it up. A separate scheduled run varies deliberately — one dimension at a time, so a failure can be attributed — and is obliged to print the exact configuration it chose and to offer a one-parameter replay. A failure from the varying run is a defect that ships with a pinned reproducing case, never a rerun. Before turning variation on, name the owner and fund the triage, because an exploratory run nobody triages becomes permanent ignored red. Keep the pinned defaults in the shared runner configuration so teams cannot drift apart, and review the exceptions rather than banning them. Then measure: defects found per triage hour decides whether the varying tier keeps its budget.

go deeper

for a junior

Be ready to say why the run that blocks a merge has to be reproducible, and why a failure caused by which machine picked up the job is not useful feedback to anyone.

for a middle

Explain the two tiers and what each is for, and the obligation on a varying run to print the configuration it chose so its failures can be replayed rather than rerun.

for a senior

Show the operating detail: one varied dimension per run, a pinned reproducing case as the exit path for every finding, and the refusal to let an exploratory run sit red.

for a principal

Own the economics and the ownership: who is funded to triage variation, where the pinned defaults live so teams cannot drift, which residual ambient risks cannot be pinned, and the measure that decides whether the varying tier keeps its budget.

## The tension Pinning and varying pull in opposite directions and both are right about something. Pinning buys reproducibility. A blocking run must be a signal: when it goes red, work stops, and the person who stopped it needs to reproduce the failure on demand. Every ambient input the run does not control is a way for that signal to lie — a failure caused by which machine picked the job up, not by the change under review. Teams that have lived through a noisy blocking suite pin everything, and they are correct to. Varying buys coverage. A fully pinned estate tests one point in the ambient input space: one seed, one zone, one locale, one encoding. Every assumption the code makes about the others survives untouched — until production, which is the varying environment by definition. Users arrive with their own locales, deployments land in their own regions, and the assumption surfaces there instead, where the feedback loop is measured in incidents rather than minutes. Neither extreme survives contact. Pinning everything everywhere is a decision to discover ambient defects in production. Varying in the blocking run is a decision to teach the team that red means nothing. ## The two-tier answer **Tier one: the gate.** Pinned completely — seed, zone, locale, encoding, and anything else ambient the estate has learned about. Deterministic by policy. If it fails, the change failed. **Tier two: exploration.** A scheduled run, not blocking, that varies deliberately. Three obligations make it useful rather than noise: *Vary one dimension at a time.* Vary the seed, or the zone, or the locale — not three at once. When a failure arrives, the varied dimension is the hypothesis, and that is most of the diagnosis. Varying several simultaneously converts a diagnosis into a search. *Emit the configuration and a replay.* Every run prints exactly what it chose and how to re-run with those values pinned. A varying run that cannot be replayed produces failures indistinguishable from infrastructure noise. *Convert every failure into a pinned case.* The exit path for a tier-two failure is a permanent case in tier one with the offending values written into it. That is what stops the same defect being rediscovered every quarter, and it grows the pinned set into a record of the ambient assumptions the code has actually broken. ## The entry condition nobody enforces Turning on a varying run is cheap; triaging it is not. An exploratory run with no named owner and no triage budget becomes permanent red within about a month, and permanent red is worse than no run at all, because it also trains people to ignore the dashboard where the real failures appear. So the honest entry condition is a named owner and funded triage time, agreed before the run is enabled — not the run first and the owner later. The measurement that keeps it honest is defects found per triage hour. If the varying tier has produced three real ambient defects in a year for a few hours of triage a month, it is one of the cheapest sources of defects in the estate. If it has produced none, either the code is genuinely portable or the run is varying the wrong dimension, and both conclusions are actionable. ## Where the defaults live Pinned defaults belong in the shared runner configuration that every team inherits, not copied into each suite. Copies drift, and drifted defaults mean a defect reproduces on one team's machines and not another's. Expect exceptions — a team whose product genuinely needs a different zone — and review them rather than forbidding them; a forbidden override becomes an undocumented one. Some things cannot be pinned at all: a third-party sandbox that returns its own dates, a device clock, an upstream system's locale. Name them explicitly as the estate's residual ambient risk and decide per case whether to bound them with a stand-in or accept the variance. ## A worked example A fleet telematics ingest ran fully pinned for a year: green gate, no ambient failures. A tier-two run varying only the locale was enabled with two hours a week of triage budget. Within a fortnight it reported an intermittent timeout in the ingest at its 1,200-request-per-minute peak — under a comma-decimal locale, a coordinate parse fell to a slower fallback path in the ingest's hot loop, and the batch exceeded its deadline about one run in nine. Note what made it usable: the run named the varied dimension and the exact value, so the hypothesis was immediate; the failure was replayable by pinning that locale; and the resolution was a permanent pinned case in tier one, at that locale, asserting the parse path. Note also what it cost: a triage budget somebody had agreed to fund before the first failure arrived.

  • Why vary only one ambient dimension per exploratory run?
    Attribution. With one varied dimension, the failure arrives with its hypothesis attached and diagnosis is usually minutes. With three, you have a failure and a search space, and the cheapest next move is to re-run varying one at a time anyway — so you have simply deferred the work while spending the triage budget.
  • What is the entry condition before enabling a varying run at all?
    A named owner and funded triage time. Variation produces failures on somebody's ordinary Tuesday, and if nobody is accountable the run goes permanently red inside a month. Permanent red is worse than no run, because it also teaches people to ignore the place where genuine failures appear.
  • How do you stop individual teams quietly overriding the pinned defaults?
    Put the defaults in the shared runner configuration every team inherits, add a build check that fails when a suite sets the process-wide values itself, and make overrides a reviewed exception rather than a prohibition. A forbidden override does not disappear; it becomes an undocumented one that nobody can reason about later.

Pinning is the lab bench and varying is the road test. A vehicle signed off on either one alone is a vehicle nobody should drive.

saying these in an interview costs you the question

  • Runs the blocking gate with a fresh configuration each time
  • Treats every varying-run failure as a flake to rerun
  • Enables variation with nobody assigned to triage it
  • Varies several ambient dimensions in one run
  • Assumes one pinned configuration proves the code portable
  • Leaves an exploratory run red indefinitely

context