skip to content

How do you choose the interaction strength for a large configuration matrix, and defend the cost?

level: principalimportance: should knowfreq 38%

answer

  1. strength is a budget, not a setting
  2. cost climbs with the widest domains
  3. not uniform across the whole model
  4. raise it where a mechanism conspires
  5. state the residual risk explicitly

basics

~20 s

Start at strength 2 everywhere because it is cheap, then raise strength only on a small risk-ranked sub-model where a mechanism makes higher-order interaction plausible. Defend the choice with run-time budget, escaped-defect evidence and stated residual risk.

solid answer

~50 s

Treat strength as a budget allocation, not a global setting. Row count grows roughly with the product of the *t* largest value counts, so moving a whole model from strength 2 to 3 is typically an order-of-magnitude jump — for a six-parameter model that can mean going from a set that fits an evening to one that cannot fit a 3-week release train. The defensible design is **mixed strength**: strength 2 as the default, strength 3 over a narrow sub-model of the parameters that touch a mechanism where three values plausibly conspire, plus seeded rows for configurations that must be evidenced. Justify the split with three inputs — the execution budget you actually have, the interaction order of defects that have escaped historically, and the blast radius of the subsystem. Then state the residual risk explicitly, review it when escapes contradict it, and keep the model versioned so the argument can be re-run rather than re-argued.

go deeper

for a junior

You are not expected to set this. Know that strength means how many parameters interact in the guarantee, that 2 is the usual default, and that higher strength means many more configurations to run.

for a middle

Be able to explain the cost curve — row count driven by the product of the largest domains at the chosen strength — and why that makes a global jump from 2 to 3 far more expensive than it sounds.

for a senior

Show you can justify a tier from evidence: reconstruct the interaction order of past escapes, identify the mechanism that makes higher-order conspiracy plausible, and keep the elevated sub-model small enough to run.

for a principal

Own the whole tradeoff: budget, escape evidence, blast radius, stated residual risk, a committed and versioned array, and a written trigger that reopens the decision when reality contradicts it.

## Why strength is a budget decision Interaction strength *t* is the promise that every combination of values across any *t* parameters appears in some row. Its cost is not linear. Covering-array size grows with the product of the *t* largest value counts and only logarithmically with the number of parameters. For the quote-engine model — value counts 4, 3, 7, 2, 3, 5 — the strength-2 floor is 7 x 5 = 35 and a generator returned 43 rows. The strength-3 floor is 7 x 5 x 4 = 140, and generated sets land well above it. At roughly 90 seconds a row, 43 rows is about an hour and fits comfortably inside the configuration slot of a 3-week release train; a few hundred rows does not, and the honest consequence of setting strength 3 globally is that the suite gets skipped rather than run. That is the frame to bring to the interview: not "what strength is correct" but "what is the execution budget, and where does spending it buy the most exposure". ## The evidence, and how much weight it carries The usual justification for strength 2 is the finding that a large majority of failures are triggered by one or two interacting parameters, with the remainder needing three or more, and that the tail beyond a handful of parameters is thin. Published studies of fielded systems do report results in that direction. But the proportions vary substantially between the domains studied, and they depend on how a "parameter" was carved out of the system in the first place — a modelling choice, not a measurement. So use the claim as a **default and a prior**, not as a proof. A principal who quotes it as a settled percentage is overclaiming; one who ignores it and demands exhaustive testing is not doing capacity planning either. The defensible position is: start at 2 because the evidence points that way and the price is low, and let your own escape data update the prior. ## Mixed strength is the actual answer Uniform strength wastes budget at both ends: it over-tests parameters that never interact and under-tests the two or three that carry a real mechanism. The design that survives review is a tiered one. *Default tier, strength 2.* Every parameter in the model. Cheap, regenerated when the model changes, committed as a versioned artefact. *Elevated tier, strength 3.* A small sub-model, three or four parameters, chosen because a mechanism makes higher-order conspiracy plausible: anything spanning a partial write and its compensation, shared caches, feature flags that gate the same code path, or the parameters implicated in a defect that already escaped. In the quote engine, the parameters touching the pricing write, the document write and the compensating rollback are the elevated tier — that is exactly where a partial-failure rollback left orphaned quotes for four consecutive release trains under a clean strength-2 suite. *Seeded rows.* Configurations that must be evidenced by name — an audited setup, the largest broker's configuration, the last two production escapes. These cost almost nothing because their pairs count toward coverage. ## The inputs you should be able to name **Execution budget.** Rows times per-row cost against the window you actually own, including environment setup and the failure-triage time a larger set generates. **Escape history.** For each defect that reached production, reconstruct the interaction order that would have caught it. If escapes are consistently order 2, your suite is not the problem and more strength is not the fix. If several are order 3 in one subsystem, that names your elevated tier. **Blast radius.** Money, safety, regulatory exposure and reversibility. A subsystem you can roll back in minutes justifies less pre-release strength than one whose failure is a mis-priced policy sold to thousands of customers. **Model quality.** Strength cannot rescue a bad parameter model. If the values per parameter are poorly chosen or the constraints are stale, spending on strength buys thoroughly covered nonsense — fix the model first, and be willing to say so when someone asks for strength 3. ## What you own as a lead Three things beyond the number. First, **stated residual risk**: write down what strength 2 does not promise, so the decision is a documented tradeoff rather than an implied guarantee. Second, **stability**: the generated array is committed and regenerated on model change, not on every pipeline run, or failure comparison across releases becomes meaningless. Third, a **review trigger**: a rule such as "any production escape whose interaction order exceeds our strength in that area reopens the tier decision", which turns the choice into something evidence updates instead of something re-argued by whoever was burned most recently. ## The anti-patterns Setting strength 3 globally to look rigorous, and quietly disabling the suite when it overruns. Rotating a random subset of the cross-product each run so nothing is guaranteed and nothing is comparable. Deferring the full cross-product to a mythical end-of-train run that never happens. Treating the covering-array percentage as a quality metric to be maximised, when the actual quality question is what each row asserts.

  • Someone asks why you do not simply set strength 3 everywhere. What is your answer?
    That the row count grows with the product of the three largest value counts rather than the two, which for a typical configuration model is roughly an order of magnitude. The suite then overruns the window it has to run in, and the real-world consequence is that it gets skipped or truncated, leaving less coverage than the cheaper design delivered reliably. I would rather guarantee strength 2 everywhere and strength 3 where a mechanism justifies it.
  • How would you decide which parameters go into the elevated-strength sub-model?
    By mechanism and by evidence, not by intuition. Mechanism: parameters that meet in the same code path — a write and its compensation, a shared cache, flags gating one branch — are where three values can plausibly conspire. Evidence: reconstruct the interaction order of past escapes and see which parameters recur. If neither test names a subset, the honest answer is that no elevated tier is justified yet.
  • What would make you lower strength rather than raise it?
    Evidence that the suite is not where the defects are. If escapes are consistently single-parameter or requirement gaps rather than interactions, extra strength adds run time and triage load without changing the outcome, and the budget belongs on the oracle or on a different technique. I would also drop strength on a subsystem being retired, and on any area where the parameter model has decayed enough that the array is covering values nobody uses.
  • How do you present this decision to people who expect every combination to be tested?
    With the arithmetic and the tradeoff, not with reassurance. Show the cross-product size, the execution cost, the window available, and what the chosen design does and does not guarantee. Then name the residual risk and the trigger that would change the decision. Stakeholders accept a bounded, written risk far more readily than a vague claim of thoroughness that collapses the first time something escapes.

saying these in an interview costs you the question

  • Picks a strength without knowing the execution budget
  • Quotes the one-or-two-parameter fault claim as settled fact
  • Applies one interaction strength uniformly across every parameter
  • Promises exhaustive coverage will happen later in the cycle
  • Treats coverage percentage as the quality target
  • Never revisits the choice when escapes contradict it

context