When is a pairwise-only configuration suite unsafe, and what do you add to close the gap?
answer
- the guarantee stops at two
- three values in one row, not three rows
- static configurations say nothing about order
- some configurations must be evidenced by name
- seeding and targeted higher strength
basics
~20 sPairwise is unsafe wherever failure needs three or more values together, depends on operation order, or must be evidenced for specific named configurations. Close the gap by seeding required rows, raising strength on a risk-ranked subset, and asserting real outcomes.
solid answer
~50 sA strength-2 array guarantees pairs and nothing else, so four situations break it. **Three-way faults**: a pair covered in two different rows is not the same as three values meeting in one row. **Order dependence**: the array says which values coexist, never in what sequence operations ran, so a fault that needs a retry after a partial-failure rollback is invisible to it. **Named obligations**: a regulator or a large customer wants evidence for their exact configuration, and "it was pairwise-covered" is not that evidence. **A weak oracle**: rows that only assert "no crash" will pass while the engine returns a wrong premium. The remedies are specific — seed the required and known-bad configurations so they are guaranteed present, raise strength to 3 on a small risk-ranked sub-model rather than globally, add sequence-based cases separately, and make each row assert a real expected outcome.
go deeper
Remember the shape of the limit: covering every pair does not mean covering every trio, and a set of configurations says nothing about the order operations happened in.
Explain why a pair appearing in two separate rows does not exercise a three-value interaction, and know that seeding lets you force specific configurations into the generated set.
Diagnose from an escape. Reconstruct which interaction order the defect needed, decide whether the array could ever have caught it, and pick a proportionate remedy rather than inflating strength across the whole model.
Own the residual-risk argument: what a strength-2 guarantee leaves uncovered, where the organisation spends more, what the oracle asserts, and how the generated set is versioned so results stay comparable across releases.
## The guarantee, stated precisely A strength-2 covering array guarantees exactly one thing: every combination of values from any two parameters occurs in some row. Everything a candidate wants to believe beyond that — that three-value interactions get "mostly" covered, that the set is representative, that unusual configurations are included — is outside the guarantee. Senior judgement here is mostly the discipline of not extending the promise. ## Failure mode one: interactions above strength 2 The classic worked example. In the quote engine, quotes for the umbrella product line, on the monthly cadence, in rating region R3 hit a **partial-failure rollback**: the pricing write succeeds, the downstream document write fails, the compensating rollback misses the pricing row, and an orphaned quote survives with no document. Any *two* of those three values are perfectly healthy. The 43-row covering array contained umbrella+monthly, umbrella+R3 and monthly+R3 — each in a different row — and passed cleanly for four consecutive 3-week release trains before a customer reported a quote with no paperwork. This is not a defect in the array; it is the guarantee working as documented. The mitigation is not to distrust pairwise but to be explicit that strength-2 buys strength-2. Where three-way exposure is plausible — anything touching partial writes, compensations, caching or feature interaction — raise strength on a narrow sub-model of the parameters actually implicated, and leave the rest at 2. Going to strength 3 across all six parameters would have taken the set from 43 rows to several hundred; strength 3 across the three parameters that touch the write path cost a couple of dozen. ## Failure mode two: order and history A covering array is a set of static configurations. It has no notion of "the second attempt after a timeout" or "renew, then endorse, then cancel". Faults that live in sequence — stale state carried between operations, a retry that double-applies a discount — are structurally outside it, and no increase in strength reaches them. That material belongs to a different design technique and should be planned as a separate body of cases rather than hopefully bolted onto the configuration run. ## Failure mode three: obligations attached to named configurations Sometimes the requirement is not statistical coverage but evidence for a specific configuration: an audit that names the combinations in scope, a large broker whose exact setup must be shown working before a release, a production configuration you have already burned once. A generated array might happen to include those rows, and might drop them on the next generation. The fix is **seeding**: supply those configurations as mandatory rows before generation. A good generator includes them verbatim, counts the pairs they already cover, and generates only what is still missing — so seeding usually *reduces* total rows while making the guarantee explicit. Seed the last few production-escaped configurations too; they cost almost nothing and they answer "would we catch that again?" with a row rather than an opinion. ## Failure mode four: the oracle, not the array The most common real-world failure is not the array at all. A configuration suite that asserts only that the run completed will happily pass 43 rows while a whole product line prices incorrectly. The array chooses inputs; something else has to decide whether each outcome was right. That may be a reference calculation, an invariant that must hold for any configuration, a comparison against the previous release's output for the same row, or explicitly stated expected values for the seeded rows. Any row whose only assertion is the absence of a crash is a row that will not repay its run time. ## Two secondary traps **Instability across runs.** If you regenerate the array on every pipeline run, a failing row may not exist next run, and week-to-week comparison collapses. Generate deliberately, commit the array, and regenerate when the model changes — treating the generated set as a versioned artefact rather than as pipeline output. **Values chosen badly.** The array reasons only over the values you supplied for each parameter. If those values are not representative of the parameter's real behaviour, a perfectly covered array tests the wrong thing thoroughly. The array's quality is capped by the model's, and the model is a human judgement call made before generation ever starts. ## How to answer the interview question Name the guarantee, name the categories it excludes, and attach a concrete remedy to each rather than concluding "so pairwise is unreliable". The strong answer is proportionate: keep strength 2 as the default because it is cheap, spend the extra strength where a mechanism makes higher-order interaction plausible, seed what must be evidenced, and put the real effort into what each row asserts.
- A pair appears in two different rows; why does that not cover the triple containing it?Because a fault of order three needs all three values present in the same execution. Covering umbrella+monthly in one row and umbrella+R3 in another means the system never ran with all three at once, so the mechanism that fails was never assembled. A strength-2 array will contain some triples incidentally, but which ones is an accident of generation, not a guarantee, and it changes when the array is regenerated.
- How does seeding a configuration change the size of the generated set?It usually shrinks it. Seeded rows are inserted first, and the pairs they already carry are marked covered, so the generator has fewer pairs left to place. Seeding therefore buys a guarantee — this exact configuration will be run — at close to zero row cost. The exception is a seeded row full of rare values that share nothing with the rest of the model, which contributes little coverage of its own.
- Would raising the whole model to strength 3 have been the right response to that rollback defect?Rarely. Strength 3 across all parameters grows the set with the product of the three largest domains, which for this model means hundreds of rows and a regression run that no longer fits the release window. The proportionate response is a narrow strength-3 sub-model over the parameters that touch the failing mechanism, plus a seeded row for the exact escaped configuration, leaving everything else at strength 2.
- How do you keep a generated array useful for comparing failures across releases?Treat it as a committed artefact rather than pipeline output. Generate it when the model changes, review the diff, and version it with the model and constraints. Then a row that failed last release either still exists or visibly left the set, and you can compare results run to run. Regenerating on every run gives you a different suite each time, and any trend analysis over it is meaningless.
saying these in an interview costs you the question
- Assumes covering all pairs also covers most triples
- Raises strength globally instead of on a risky subset
- Expects a configuration array to catch order-dependent faults
- Asserts only that the run finished without crashing
- Regenerates the array every pipeline run and compares failures
- Treats generated coverage as evidence for a named configuration