skip to content

As a build matrix accumulates exclusions, when should you model it as a sum of disjoint job families rather than one product with exceptions?

level: principalimportance: should knowfreq 33%

answer

  1. same jobs, two descriptions
  2. totals must agree after refactor
  3. exceptions interact, families do not
  4. disjointness is the licence to add
  5. distinct identities equal summed counts

basics

~20 s

Switch when the exceptions stop being a short, independent list. One product minus exceptions stays compact while exclusions are few; once exclusions interact, a sum of disjoint families keeps each part a plain product and makes the total verifiable.

solid answer

~60 s

Both models describe the same set of jobs, so they must agree on the total — the choice is about which one stays reviewable as the pipeline grows. A **single product with an exclusion list** is compact, and adding an axis multiplies through everywhere at once; its weakness is that each new exclusion makes the real count harder to state, and two exclusions that overlap subtract the same combination twice without warning. A **sum of disjoint families** makes each family a clean product, so the total is a sum you can check family by family, and a reviewer can see what actually runs; its costs are duplication and drift, since an axis change must now be repeated in every family. The switch is justified when exclusions carve the space into groups that genuinely differ — a legacy runtime supporting one architecture is really its own family — and the invariant to hold onto is that the number of distinct job identities must equal the sum of the family counts.

go deeper

for a junior

Know that a matrix can be written as one product with exceptions or as several smaller matrices added together, and that both describe the same jobs.

for a middle

Explain why the addition rule needs the families to be disjoint, and compute the total both ways on a small matrix to show the two descriptions agree.

for a senior

Detect the overlap in practice: canonical job identities, distinct count against the summed family counts, and a before-and-after comparison whenever the configuration is restructured.

for a principal

Choose the model by which failure the organisation can afford — a wrong number under a compact product, or a job silently built twice under drifting families — and record the reasoning so it is not re-decided blindly.

## Two descriptions of the same set A build matrix can be written two ways, and this is a modelling decision rather than a mathematical one. - **One product with exceptions.** Declare the axes, take the full product, then list the combinations policy removes. The count is `product - excluded`, which is complementary counting. - **A sum of disjoint families.** Declare several smaller matrices, each with its own axes, arranged so no job appears in two. The count is the sum of the per-family products, which is the addition rule. They describe the same jobs, so a refactor from one to the other must leave the total unchanged. If the number moves, the refactor changed the pipeline, and that is worth knowing before anyone celebrates a tidier configuration. ## What each model buys and costs | | One product with exceptions | Sum of disjoint families | |---|---|---| | Adding an axis | one edit, multiplies everywhere | repeated in every family | | Stating the true count | product minus a list that may interact | add the per-family products | | Failure mode | two exclusions removing the same combination twice | a job that lands in two families and runs twice | | Reviewability | must hold the whole product in your head, minus exceptions | each family readable on its own | | Drift risk | low, one source of truth | high, families diverge over time | The asymmetry worth naming: the first model's failure is a **wrong number** with the right jobs running, while the second model's failure is a **wrong pipeline** — a job built twice, or a gap nobody notices. Neither is strictly safer; they fail in different currencies. ## How to decide 1. **Count the exclusions and ask whether they interact.** A handful of independent exclusions is exactly what complementary counting handles well, and splitting into families for them buys nothing but duplication. 2. **Ask whether the exclusions have a reason in common.** Exclusions that all follow from one fact — this runtime generation supports one architecture — are describing a family, and the configuration should say so. 3. **Ask who reads it.** If the matrix is the artefact a reviewer uses to decide whether coverage is adequate, families win, because a reviewer can check one family at a time. 4. **Ask how the axes will grow.** If new axes arrive often and must apply uniformly, the single product resists drift far better. 5. **Ask whether the count must be auditable.** Budget, coverage claims and feedback-time targets are all built on the job count, and a number nobody can reproduce is a poor foundation for any of them. ## The invariant that keeps families honest The addition rule is exact only for disjoint cases, and disjointness is precisely what a family split can silently lose. Make the check mechanical: - Give every job a **canonical identity** from its axis values. - Collect the identities produced by every family. - **The number of distinct identities must equal the sum of the family counts.** If the distinct count is lower, some job belongs to two families and is being both double-counted and, worse, double-built. If a family's own jobs are fewer than its product suggests, that family has an internal exclusion of its own, which is a hint that it wants splitting again. This check also gives the refactor its safety net: run it before and after, and the totals must match. ## The trade-off a lead actually owns The decision is rarely about which arithmetic is prettier. It is about which mistake the organisation can afford. - A team that ships a new axis every month and rarely audits coverage is better served by one product with a short exclusion list, because uniform change is cheap and drift is the expensive failure. - A team whose pipeline budget is scrutinised, or whose matrix encodes support commitments that differ by platform generation, is better served by families, because every claim about what is covered can be checked against one readable block. - A team with **both** pressures usually lands on a hybrid: one product for the mainstream axes, plus a small number of explicitly named families for the genuinely different cases. That is fine, provided the disjointness invariant is enforced across the whole thing rather than within each half. And the discipline that makes either model survive: whichever you pick, write down why. The count is an argument, and an argument nobody can follow will be re-derived by the next person from scratch, differently.

  • What does the disjoint-families model cost you?
    Duplication and drift. An axis change must be repeated in every family, and families diverge as people edit the one they work on. Its characteristic failure is also worse in kind: a job landing in two families is built twice, not merely counted twice.
  • How do you check that job families really are disjoint?
    Derive a canonical identity for every job from its axis values, collect them across all families, and require the number of distinct identities to equal the sum of the family counts. A shortfall names the jobs that belong to two families.
  • Should the total change when you refactor a product with exclusions into families?
    No. Both models describe the same set of jobs, so the counts must agree. A total that moves means the refactor changed what runs, which should be a deliberate decision rather than a side effect of tidying the configuration.

saying these in an interview costs you the question

  • Splits into families without checking no job lands in two
  • Expects the two models to give different totals
  • Keeps one product and a growing exclusion list purely for brevity
  • Assumes the pipeline's reported number needs no independent count
  • Treats overlapping exclusion rules as harmless because each looks correct